Methods
diversify-text uses a pluggable method architecture. Each method is a
DiversificationMethod subclass that generates
paraphrases using a different model or algorithm.
Overview
Method |
Model Size |
Speed |
Performance |
Description |
|---|---|---|---|---|
|
~800M params |
TBD |
TBD |
Few-shot style transfer using authorship embeddings |
|
~3B params (default) |
TBD |
TBD |
Prompt-based paraphrasing using a causal LM |
|
~3B params (default) |
TBD |
TBD |
Styles defined by rewrite instructions, via a causal LM |
TinyStyler
TinyStyler is a T5-based model that performs few-shot text style transfer by conditioning on authorship-embedding representations.
Given a source text and a set of style example sentences, TinyStyler generates
a paraphrase that preserves the content while shifting toward the demonstrated
writing style. diversify-text cycles through different style groups from a
configurable style bank to produce multiple stylistically diverse outputs.
Note
TinyStyler is based on CISR style embeddings, which have been shown to work well for social-media-like settings and formality transfer. The model may not perform as expected when reproducing other styles.
Style bank. TinyStyler does not use the package’s default style
bank: it has its own small bank of styles it demonstrably handles
(informal, formal, question, …), selected by running the
model on every available style and comparing outputs
(evaluations/tinystyler_style_ratings.txt in the repository).
A few additional styles work but are more likely to produce swearing;
they are selectable by name only. The Styles page lists all of
them, along with the default bank used by the prompting method.
Citation:
@inproceedings{horvitz-etal-2024-tinystyler,
title = "{T}iny{S}tyler: Efficient Few-Shot Text Style Transfer with Authorship Embeddings",
author = "Horvitz, Zachary and
Patel, Ajay and
Singh, Kanishk and
Callison-Burch, Chris and
McKeown, Kathleen and
Yu, Zhou",
editor = "Al-Onaizan, Yaser and
Bansal, Mohit and
Chen, Yun-Nung",
booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2024",
month = nov,
year = "2024",
address = "Miami, Florida, USA",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2024.findings-emnlp.781",
pages = "13376--13390",
}
Prompting
The prompting method generates paraphrases by sending input texts to a
local HuggingFace causal language model with a prompt template. The default
model is SmolLM3-3B
using insights from The Synthetic Data Playbook.
results = diversify("The cat sat on the mat.", method="prompting")
Choosing a model. Any HuggingFace causal LM can be used. Pass the model
identifier directly to diversify():
results = diversify(
"The cat sat on the mat.",
method="prompting",
model="Qwen/Qwen3-4B-Instruct-2507",
)
or to the method constructor:
from diversify_text import Diversifier
from diversify_text.method.prompting import PromptingMethod
method = PromptingMethod(model="mistralai/Mistral-7B-Instruct-v0.3")
results = Diversifier(method=method).diversify("The cat sat on the mat.")
model works for the prompting and zero_shot methods; the
tinystyler method has a fixed model, so passing model with it
raises an error.
Instruct-tuned models are recommended. Chat templates are applied automatically when the tokenizer provides one.
Note
Thinking/reasoning models (e.g. SmolLM3-3B) are detected automatically and
have their thinking mode turned off (enable_thinking=False) during
generation. Thinking tokens add overhead without improving paraphrase
quality in this setting.
Inference backend. The method currently uses the transformers library
for inference.
Note
vLLM support, batched inference, and streaming from large files are planned for a future release.
Prompt templates. All templates are example-based style transfer prompts:
the target style is demonstrated through example texts inserted into the
prompt; prompts without style examples are intentionally not supported. The
default template is style_transfer; humanize_transfer (inspired by
Zhang et al. (2024)) additionally
instructs the model to imitate human imperfections found in the style
examples. Select a template — or pass your own — via the prompt option:
results = diversify(
"The experiment was conducted in a controlled lab setting.",
method="prompting",
method_kwargs={"prompt": "humanize_transfer"},
)
A custom template must contain both the [DOCUMENT SEGMENT] and
[STYLE EXAMPLES] placeholders ([STYLE NAME] is optional):
my_prompt = (
"Study these examples:\n[STYLE EXAMPLES]\n"
"Rewrite the following text in the same style. "
"Text: [DOCUMENT SEGMENT]"
)
results = diversify(
"The cat sat on the mat.",
method="prompting",
method_kwargs={"prompt": my_prompt},
)
Style examples. The prompting method uses the default style bank,
loaded from stylebank.json: styles organized in a
language-variation taxonomy — individual styles (idiolects) and
group-level variation across time (diachronic), region (diatopic),
social group (diastratic), register (diaphasic), and medium
(diamesic). See diversify_text.styles.DEFAULT_STYLE_BANK and
the Styles page. Select styles with the top-level styles
parameter (or pass your own via style_texts):
results = diversify(
"The experiment was conducted in a controlled lab setting.",
method="prompting",
styles=["informational"],
)
Zero-shot
The zero_shot method defines each style by a rewrite instruction
instead of example texts, and sends one instruction per style to a
causal language model (same default model and options — model,
precision — as the prompting method).
Its own style bank maps style names to instructions
(ZeroShotMethod.style_bank:
formal, simple, complex, caps, lowercase, and
more), so styles and n select from these:
results = diversify(
"The experiment was conducted in a controlled lab setting.",
method="zero_shot",
styles=["formal", "caps"],
)
With this method, style_texts are instructions — exactly one per
style. An instruction can place the input text itself with
[DOCUMENT SEGMENT]; otherwise the text is appended at the end:
results = diversify(
"The experiment was conducted in a controlled lab setting.",
method="zero_shot",
style_texts={"pirate": ["Rewrite the text as an old-timey pirate would say it."]},
)
Development
To see the exact prompts sent to the model, enable debug logging:
import logging
logging.basicConfig(level=logging.DEBUG)
Adding a new method
See Creating a custom method in the Usage Guide for instructions on
implementing your own DiversificationMethod.