Methodology

What Clarity measures

Clarity screens observable language functions in a pasted message—patterns that may compress informed, voluntary choice. Examples include guilt leverage, forced urgency, isolation framing, conditional access, and similar pressure functions described in the taxonomy below.

Clarity does not measure intent, personality, diagnosis, truth, relationship safety, or whether a person is “manipulative.” It highlights language that may narrow reflection space. Context, history, and tone always matter more than any score.

Language-function taxonomy (screening v0.3.2)

Twelve pattern families are evaluated in the current Controlled Beta:

When multiple functions appear together, a small cluster bonus may raise the overall screening score. Response prompts are generated per detected function.

How the experimental score is constructed

Each detected function contributes weighted points based on severity tags (mild, clear, strong). The overall score maps to three bands:

Scores are screening outputs, not probability estimates or clinical measures. They summarize matched language in the pasted text only.

Minimum input and abstention

If a message has fewer than 15 recognizable words, Clarity abstains: no numeric score is shown and the interface explains that more text is needed. Safety notices may still appear on shorter messages when explicit harm or coercion language is present.

Safety notices (separate from scoring)

Safety classification runs before pattern scoring. When language about violence, weapons, confinement, self-harm coercion, or contextual stalking is detected, Clarity shows a safety notice instead of a manipulation score. This is not a safety assessment and cannot determine whether you are in danger.

Stalking-related language uses a two-stage model: intrusive behaviors paired with menace, severe intrusion, or multiple eligible behaviors. Explicit future threats and continuous present-tense stalking with refusal phrases are handled separately.

Context handling

Attribution and context rules are applied locally to quotations and clauses:

Current version and evaluation status

Screening version: v0.3.2 (methodology and safety logic deployed with this site build).

Specificity tiers: Phrase families are classified as high, medium, or low specificity. A single low-specificity match (for example a bare “you need to” in an autonomy-supporting sentence) stays in the Low band. Moderate results require a high-specificity pattern, multiple independent pressure functions, or several clear repetitions within one medium-specificity function. High results require multiple independent functions, defined leverage clusters, or several clear pressure signals together.

Regression testing: An expanded automated regression suite (targeted audit fixtures, benign adversarial controls, and safety probes) runs before each release. This is a regression count, not a representative validation study. Sensitivity, specificity, fairness, and real-world performance on natural messages have not been established.

No representative accuracy percentage should be inferred from this count.

Known limitations and failure modes

Corrections and feedback

If you believe output is incorrect or harmful, see Contact and Corrections. Do not submit private message content when reporting issues.