Methodology

What Clarity measures

Clarity screens observable language functions in a pasted message, patterns that may compress informed, voluntary choice. Examples include guilt leverage, forced urgency, isolation framing, conditional access, and similar pressure functions described in the taxonomy below.

Clarity does not measure intent, personality, diagnosis, truth, relationship safety, or whether a person is “manipulative.” It highlights language that may narrow reflection space. Context, history, and tone always matter more than any score.

Language-function taxonomy (screening v0.3.3)

Twelve pattern families are evaluated in the current Controlled Beta:

When multiple functions appear together, a small cluster bonus may raise the overall screening score. Response prompts are generated per detected function.

How the experimental score is constructed

Each detected function contributes weighted points based on severity tags (mild, clear, strong). The overall score maps to three bands:

Scores are screening outputs, not probability estimates or clinical measures. They summarize matched language in the pasted text only.

Minimum input and abstention

If a message has fewer than 15 recognizable words, Clarity abstains: no numeric score is shown and the interface explains that more text is needed. Safety notices may still appear on shorter messages when explicit harm or coercion language is present.

Safety notices (separate from scoring)

Safety classification runs before pattern scoring. When language about violence, weapons, confinement, self-harm coercion, or contextual stalking is detected, Clarity shows a safety notice instead of a manipulation score. This is not a safety assessment and cannot determine whether you are in danger.

Stalking-related language uses a two-stage model: intrusive behaviors paired with menace, severe intrusion, or multiple eligible behaviors. Explicit future threats and continuous present-tense stalking with refusal phrases are handled separately.

Context handling

Attribution and context rules are applied locally to quotations and clauses:

Media Lens influence graph (experimental, no score)

Media Lens is a separate, unreleased preview mode for public articles, advertisements, speeches, and campaign material. It is not part of the Clarity analyzer above, does not share Clarity's code path, and is not deployed to this site. Media Lens produces no overall 0-100 score for an article, and no field labels a person or outlet as manipulative.

Media Lens separates results into four dimensions shown independently: Language (span-tied observations of influence patterns such as loaded or moralized wording, fear or threat framing, urgency, false dilemma, identity framing, scapegoating, certainty beyond evidence, vague authority, anecdote generalized into a trend, bandwagon appeals, and adversarial framing), Claims (checkable-looking statements, with support always marked "not checked" in this preview; no fact-checking path exists yet), Coverage (origin, independent reporting, and duplication, when provenance data is available), and Source context (publisher metadata shown separately, never fused into a judgment).

Every Language and Coverage observation is tied to a specific span of text or is explicitly marked unlocalized; quoted or attributed language is never presented as the author's own words. Where evidence is insufficient, Media Lens abstains rather than guessing, and abstention is a first-class result, not an error state.

Media Lens uses a typed semantic classifier called Jev, made by TypeSafe, to answer narrow, structured questions about a span of text (for example, which influence pattern it most resembles). Jev is not treated as an authority on intent, truth, outlet quality, or any person's character; deterministic rules in Media Lens's own code decide what, if anything, is shown, and a mismatched or failing classifier produces "needs review" or abstention rather than a fabricated result. Coverage data, when present, comes from Newsjack run artifacts an operator has already produced locally; Media Lens does not run a live news search.

Media Lens results are regression-tested against a fixed set of invented example articles, the same "regression count, not a representative validation study" caveat that applies to Clarity above. No accuracy, bias-detection, or fact-checking claim is made for Media Lens output.

Current version and evaluation status

Screening version: v0.3.3 (methodology and safety logic deployed with this site build).

Specificity tiers: Phrase families are classified as high, medium, or low specificity. A single low-specificity match (for example a bare “you need to” in an autonomy-supporting sentence) stays in the Low band. Moderate results require a high-specificity pattern, multiple independent pressure functions, or several clear repetitions within one medium-specificity function. High results require multiple independent functions, defined leverage clusters, or several clear pressure signals together.

Regression testing: An expanded automated regression suite (targeted audit fixtures, benign adversarial controls, safety probes, and a labeled gold harness) runs before each release. This is a regression count, not a representative validation study. Sensitivity, specificity, fairness, and real-world performance on natural messages have not been established.

No representative accuracy percentage should be inferred from this count. The gold harness fails a release if locked benign cases leave Low or locked explicit-harm cases are missed.

Known limitations and failure modes

Corrections and feedback

If you believe output is incorrect or harmful, see Contact and Corrections. Do not submit private message content when reporting issues.