AI Seer’s Facticity Harness Scores 98% on Legal Citations Benchmark, Matching Frontier Models at Up to 115 Times Lower Cost

On a 90-instance test built from real U.S. federal-court opinions, the citation-checking pipeline inside Facticity.AI correctly classified 88 cases while resolving every citation itself for $0.10 per 100 checks, making routine verification affordable for sole legal practitioners and firms alike. A preview of this was announced at a recent Tech in Asia Workshop on How to Stop AI from Giving Wrong Answers to Customers.

AI Seer’s Facticity Harness Scores 98% on Legal Citations Benchmark, Matching Frontier Models at Up to 115 Times Lower Cost
Singapore, Singapore, October 07, 2026 --(PR.com)-- AI Seer Pte. Ltd. (“AI Seer”), backed by venture capitalist Tim Draper and whose Facticity.AI was named one of TIME’s Best Inventions of 2024 is building on its strengths of being more than 3X less incorrect than other leading AI Search at news citations. AI Seer's Founder reiterated this strength, but also previewed at a full and well-received Tech in Asia Workshop on Facticity.AI's unique edge in not being sychophantic and how even frontier models still continue to produce the right claim but the wrong source. AI Seer also published results from its Hallucinated Legal Citations Benchmarking. a 90-instance test of whether an AI system can tell when a legal citation actually supports the claim attached to it. The Facticity harness, the citation-checking pipeline deployed inside AI Seer’s Facticity.AI verification engine, correctly classified 88 of the 90 instances (98%) at a metered cost of $0.10 per 100 checks. Nine one-shot frontier models were run on the same instances; those that matched or exceeded the harness’s accuracy cost between 17 and 115 times more.

The problem of secondary hallucinations shows up particularly prominently in law the checks and balances with higher courts and opposing counsel also make the problem particularly visible with more than 2100+ cases (many recent in the AI Hallucination Cases Database). However, its present in all knowledge work, and AI Seer is in the midst of scaling up a paid paid pilot for its API with a Big Four Consulting firm. The risk is not confined to general-purpose chatbots used in our internal benchmarking: A 2024 Stanford study found that AI research tools from major legal publishers hallucinated in at least one in six queries. Grace Chong, Head of Financial Regulatory practice at Drew & Napier LLC says that “Facticity delivers clear, detailed responses promptly, supported by an extensive range of sources and relevant extracts. As lawyers, we need to understand the authority and reasoning behind a conclusion, as well as the conclusion itself. This has made it easier for us to trace the analysis to its underlying sources, assess their relevance and verify the accuracy of the response, while reducing the time required for independent cross-checking. That transparency is particularly valuable in legal and regulatory research, where precision and context matter.” Legal professionals can try it out for themselves at app.facticity.ai/citation-check.

Founder Dennis Yap, noted that a frontier-model plan does not mean every citation check should use the most expensive model. Facticity can be right-sized for Enterprise AI, AI workspaces and Agentic Systems through its different integrations or customized for you. The results matter most to the lawyers with the least cover: small firms, sole practitioners, or freelance lawyers who draft with generative AI under time pressure, without a knowledge-management team or an enterprise verification stack behind them. Even as recently as 2026, Prososki v. Regan, 321 Neb. 38 (2026); Nebraska Public Media, "Nebraska Supreme Court blasts AI-authored court filings" 90% of the citations were wrong, and the lawyer got suspended indefinitely. A large firm that adopts generative AI can afford a review layer: a second associate, a knowledge-management team, an enterprise licence for a verification tool. A sole practitioner or a three-lawyer firm usually cannot, and it is the same practitioner who is most likely to be drafting at midnight with a chatbot open in the next tab. The sanctions that follow a hallucinated citation - a fine, a referral to the bar, a suspension, a client who reads about it in the news - land on a practice with no margin to absorb them. AI Seer's Hallucinated Citations Benchmark contains 90 citation-verification instances derived from real U.S. federal-court opinions in CLERC, the Case Law Evaluation and Retrieval Corpus. Thirty genuine citation examples were each transformed into three test classes: Good, Irrelevant and Reversal. For greater detail and context, please read this article.
Contact
AI Seer Pte. Ltd.
Karina Caunane
65 89007408
app.facticity.ai/citation-check
Please contact through LI (https://www.linkedin.com/in/karina-caunane-0706aa40/) before trying to call.
ContactContact
Multimedia
Citation-verification accuracy versus cost.

Citation-verification accuracy versus cost.

The top systems cluster within two accuracy points; their costs span more than 100×. Source: AI Seer, Hallucinated Citations Benchmark. https://github.com/Reality-Detector/hallucinated-citations-benchmark

The Tech in Asia Workshop Crowd

The Tech in Asia Workshop Crowd

Dennis was the opening speaker of the Implementation Stage. See more: https://x.com/ArAIstotle/status/2099753767843549453/photo/3

Dennis talking about rightsizing Facticity.AI Verification for your AI workflows

Dennis talking about rightsizing Facticity.AI Verification for your AI workflows

@facticityai / ArAIstotle was brought into the world of AI workflow integrations through ACP / MCP, which is just a step away from smaller enterprises! enterprises. See more: https://x.com/ArAIstotle/status/2099753767843549453?

Categories

Sign In

New to PR.com?

Get email news alerts fitting your preferences