Enterprise

Independent AI Observatory Reveals Usage Gaps in Company Reports

Stanford and MIT researchers analyzed 24,000 real conversations to show how AI labs filter out sensitive content from public data.

Omega Editorial· August 18, 2026· 3 min read

Independent researchers challenge AI company transparency

A new research initiative is exposing significant blind spots in how AI companies report user behavior. The AI Observatory, led by researchers from Stanford and MIT, analyzed more than 24,000 real conversations across ChatGPT, Claude, Gemini, and other models—revealing that corporate usage reports systematically exclude large portions of actual user activity.

According to MIT Technology Review, which first reported the findings, the project was created specifically to provide independent verification of claims made by companies like Anthropic and OpenAI. "There is no independent source to corroborate it," says Anka Reuel, a Stanford PhD candidate who co-leads the research.

What company reports leave out

The Observatory's analysis found stark differences between public company data and real-world usage patterns. When researchers applied Anthropic's own filtering methodology to their dataset, 48% of conversations would have been excluded from the company's Economic Index reports.

Those filtered conversations were substantially more likely to involve health and relationships (44.2% versus 31.2% in Anthropic's published analysis), adult or illicit topics (7.9% versus 2.1%), harassment and hate speech (27.5% versus 5.66%), and sexual content (16.7% versus 2.4%).

Anthropic's Economic Index explicitly focuses on work and productivity uses of Claude, filtering out unrelated conversations. While the company has published separate analyses of companionship use and concerning content like CSAM generation, independent researchers say the fragmented approach obscures the full picture.

Usage varies dramatically by model and time

The Observatory documented significant differences in how people interact with different AI platforms. Users turned to Grok and Gemini primarily for information retrieval, with Grok particularly popular for news and politics—but also where misinformation concentrated most heavily. People used Claude more for coding tasks, Gemini for social roleplay, and ChatGPT for homework help.

Even different versions of the same model showed distinct patterns. Conversations with GPT-3.5 were notably shorter than those with GPT-4o, which became associated with longer, more iterative exchanges and reports of emotional dependency.

Over the 2023-2025 period the datasets covered, conversations generally grew longer and more elaborate. Small talk increased while AI self-disclosure as chatbots decreased, suggesting rising use for companionship. Sensitive content exchanges dropped, potentially indicating more effective platform safeguards.

Why it matters

Policymakers and researchers are making consequential decisions about AI regulation and deployment based on incomplete information controlled entirely by the companies being regulated. Without independent verification, there's no way to assess whether general-purpose AI systems are used "mostly for good, or mostly for bad," as UT-Austin assistant professor David Widder notes. The AI Observatory's work provides a crucial check on corporate narratives, though its 24,521 conversations remain a fraction of the millions of exchanges the major labs analyze internally.

Dataset limitations and future plans

The Observatory aggregated conversations from seven existing datasets collected with user consent, covering 85,633 conversational turns across 52 different models from 5,000 users. The researchers acknowledge their voluntary data sources likely underrepresent sensitive uses that people hesitate to share.

The team plans to expand the dataset over time and make it available to other researchers. Reuel says the ideal solution would be for AI companies to share their data with independent researchers in privacy-preserving ways—but until then, decision-makers risk "operating in the wild" without understanding what's actually happening beyond company narratives.

These findings were first reported by MIT Technology Review.

#ai transparency#ai safety#chatgpt#claude#ai research#content moderation

This is an original analysis by the Omega editorial team. Source reporting: AI Watch.

Want systems like this working for your business?

Book a Call

More in Enterprise

Enterprise· 3 min read

Walmart workers correct AI tools they're expected to use daily

The retailer's massive AI rollout reveals the friction when frontline employees become de facto trainers for workplace automation.

Via AI Watch · Aug 18, 2026
Enterprise· 3 min read

Marketing agencies adopt AI without fixed guardrails

Industry leaders embrace fluid frameworks as entry-level workers face new barriers to traditional career paths.

Via AI Watch · Aug 18, 2026
Enterprise· 4 min read

B2B Marketing Automation Platforms Split Into Two Camps

HubSpot's breadth versus Pardot's Salesforce depth creates divergent paths with real cost implications beyond subscription fees.

Via Automation Watch · Aug 18, 2026