Methodology
Last updated: August 3, 2026
What Clarity measures
Clarity screens observable language functions in a pasted message—patterns that may compress informed, voluntary choice. Examples include guilt leverage, forced urgency, isolation framing, conditional access, and similar pressure functions described in the taxonomy below.
Clarity does not measure intent, personality, diagnosis, truth, relationship safety, or whether a person is “manipulative.” It highlights language that may narrow reflection space. Context, history, and tone always matter more than any score.
Language-function taxonomy (screening v0.3.2)
Twelve pattern families are evaluated in the current Controlled Beta:
- Guilt leverage — care or loyalty used as currency for compliance.
- Forced urgency — compressed decision windows.
- Conditional threat — cooperation linked to penalties.
- Isolation — secrecy or discouraging outside input.
- Dismissal — feelings or concerns minimized.
- Misleading absolutes — rigid always/never framing.
- Assigned obligation — duties imposed without mutual agreement.
- Minimization — impact of concerns downplayed.
- Resigned withdrawal — cooperation withdrawn to create pressure.
- Implied rejection — connection threatened indirectly.
- Conditional access — affection, contact, or access withheld for compliance.
- Responsibility shift — blame placed on the recipient for the sender’s state.
When multiple functions appear together, a small cluster bonus may raise the overall screening score. Response prompts are generated per detected function.
How the experimental score is constructed
Each detected function contributes weighted points based on severity tags (mild, clear, strong). The overall score maps to three bands:
- Low (0–30): few or mild patterns.
- Moderate (31–60): clear functions that may narrow reflection space.
- High (61–100): multiple strong pressure functions together.
Scores are screening outputs, not probability estimates or clinical measures. They summarize matched language in the pasted text only.
Minimum input and abstention
If a message has fewer than 15 recognizable words, Clarity abstains: no numeric score is shown and the interface explains that more text is needed. Safety notices may still appear on shorter messages when explicit harm or coercion language is present.
Safety notices (separate from scoring)
Safety classification runs before pattern scoring. When language about violence, weapons, confinement, self-harm coercion, or contextual stalking is detected, Clarity shows a safety notice instead of a manipulation score. This is not a safety assessment and cannot determine whether you are in danger.
Stalking-related language uses a two-stage model: intrusive behaviors paired with menace, severe intrusion, or multiple eligible behaviors. Explicit future threats and continuous present-tense stalking with refusal phrases are handled separately.
Context handling
Attribution and context rules are applied locally to quotations and clauses:
- Training, fiction, and news quotations — attributed violence or stalking language in examples is excluded from safety escalation.
- Mixed messages — a real first-person threat in a separate sentence or clause after an attributed quote still triggers a safety notice.
- Negation — local negation (e.g. “would never hurt you”) can exclude a matched phrase unless a later clause contains a real threat.
- Consent and invitations — benign consent phrases (invitations, requested deliveries, permission boundaries) reduce false positives on location language.
- Refusal violations — phrases such as “told me to stop,” “asked me to leave,” or “told me to go away” pair with stalking behaviors to trigger notices.
Current version and evaluation status
Screening version: v0.3.2 (methodology and safety logic deployed with this site build).
Specificity tiers: Phrase families are classified as high, medium, or low specificity. A single low-specificity match (for example a bare “you need to” in an autonomy-supporting sentence) stays in the Low band. Moderate results require a high-specificity pattern, multiple independent pressure functions, or several clear repetitions within one medium-specificity function. High results require multiple independent functions, defined leverage clusters, or several clear pressure signals together.
Regression testing: An expanded automated regression suite (targeted audit fixtures, benign adversarial controls, and safety probes) runs before each release. This is a regression count, not a representative validation study. Sensitivity, specificity, fairness, and real-world performance on natural messages have not been established.
No representative accuracy percentage should be inferred from this count.
Known limitations and failure modes
- Fixed pattern lists miss many forms of pressure and may flag benign language.
- Screenshot OCR can introduce noise that affects word counts and matches.
- Irony, sarcasm, cultural context, and relationship history are not modeled.
- Safety rules prioritize catching explicit harm language; false negatives remain possible.
- Scores apply to pasted English text. Multi-message threads show per-message results; the overall band reflects the highest-scoring eligible message in the thread.
- Image upload is disabled in this Beta candidate. OCR is English-only when enabled in future releases.
Corrections and feedback
If you believe output is incorrect or harmful, see Contact and Corrections. Do not submit private message content when reporting issues.