Open source
Open-source research toolFree browser app2026

CiteScope

A citation is not the same thing as evidence.

I built CiteScope to make that distinction easier to investigate. Bring an AI-generated answer and its cited source text, inspect what the references actually contain, and keep a review record you can revisit.

C I T E S C O P E   /   E V I D E N C E   R E V I E WIllustrative example, not a measured result

AI-GENERATED CLAIM

Retrying an API request always guarantees success.[2]

Reference [2] exists. But does the source support the statement?

CITED SOURCE [2]

Contradiction to review
Retrying an unsuccessful API call does not guarantee it will succeed.

CiteScope verifies the reference structure. A reviewer decides whether this passage contradicts the claim.

Structural checks are automatic. Evidence judgments are not.

What it helps you do

The public version is a practical evidence-review workspace, not a one-click factual-truth detector.

01

Inspect the references

Paste an AI-generated answer and the numbered source snapshots it cites. CiteScope finds missing source IDs and checks whether quoted passages appear in the cited text.

02

Review the actual claim

Edit or split statements into smaller, reviewable claims. A person can then label cited evidence as supportive, partial, unsupported or contradictory.

03

Keep the evidence

Export a structured dataset or a readable report, and import an earlier review to continue. The Python toolkit can compare annotated results across model runs.

Under the hood

CiteScope extends the published GEO research project, with the original authors, source and Apache-2.0 license credited. My additions include a browser-based review interface, Python citation auditing, source-variant experiments, benchmark tooling and optional local language-model suggestions.

Its browser app runs without accounts or paid API calls. The optional Hugging Face NLI model runs separately and provides suggestions; it is not automatic fact-checking in the hosted app.

The research benchmark is still being developed. Current example results are demonstrations, not independently validated claims about AI model accuracy or search ranking.

Explore the project

Explore next

Have something worth fixing?Start a conversation