Thomson Reuters Builds Proprietary Legal AI Model, Posts 0.83 Factuality Score
The company trained its own large language model on decades of legal content rather than wrapping a general-purpose AI, targeting accuracy over breadth.

Thomson Reuters takes different path in legal AI race
Thomson Reuters has launched its own large language model for legal work rather than building on top of general-purpose AI systems like most competitors. The model, called Thomson, was trained on decades of proprietary content from Westlaw, Practical Law, Checkpoint, and Reuters, with validation from the company's subject matter experts.
The approach differs from most legal AI products currently on the market, which layer interfaces and features over existing foundation models from OpenAI, Anthropic, or Google. Thomson Reuters started with an open-source foundation but then trained it specifically on legal content and reasoning, according to details first reported by Thomson Reuters Legal.
Performance centered on verifiable accuracy
In internal evaluations, Thomson scored 0.83 on factuality when working with Westlaw and Practical Law content. That metric measures whether claims made by the model can be traced back to sources that actually support them. By comparison, leading frontier models with open web access scored between 0.65 and 0.68 on the same measure.
The company also tested Thomson on PrBench Legal Hard, a demanding legal reasoning benchmark, where it posted the top score among models evaluated. Two independent academics who tested the model reached similar conclusions about its citation quality and response accuracy.
Professor Samuel Dahan of Queen's Conflict Analytics Lab and Cornell Legal AI Lab found Thomson's citation quality competitive with leading models, even on Canadian employment law questions outside its primary training focus. Professor Jonathan H. Choi of Washington University School of Law tested it against corporate tax questions and noted a preference for Thomson's responses, particularly citing its references to underlying treatises.
Deployment strategy and training approach
Thomson is now deployed in Tabular Analysis within CoCounsel Legal, handling structured document review tasks. The company maintains a multi-model approach, applying Thomson where it delivers clear advantages while using other leading models elsewhere.
Partner-level practitioners built evaluation rubrics for complex research questions during training. The company reports using less than 10 percent of its proprietary content library so far, indicating substantial room for model refinement.
Thomson Reuters emphasized that the model is not trained on customer data and will not be without explicit consent. The model runs on infrastructure the company controls, addressing data governance concerns common in legal departments evaluating AI tools.
Why it matters
Legal departments face mounting pressure to adopt AI tools, but most products on the market share the same underlying technology with different interfaces. A purpose-built model trained on authoritative legal content and validated by practitioners represents a fundamentally different approach to accuracy and reliability. For organizations with fiduciary duties, the distinction between a general-purpose model optimized for breadth and a specialized model optimized for verifiable legal reasoning carries direct risk implications. The factuality gap between Thomson and general models—0.83 versus 0.65-0.68—quantifies a difference that matters when errors carry professional liability.
These details were first reported by Thomson Reuters Legal in a company blog post announcing the model launch.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call