Another member of the Big Four has been forced to answer uncomfortable questions about its use of generative AI.
Researchers have found that several thought leadership reports published by PwC Middle East contained fabricated citations, broken references, unverifiable claims, and at least one academic paper that appears never to have existed. The findings mirror similar incidents involving KPMG and EY earlier this year, suggesting the issue is no longer an isolated editorial mistake but part of a broader challenge facing organizations that increasingly rely on AI-assisted content production without sufficiently rigorous human verification.
The irony is difficult to ignore. These are the same firms advising governments and multinational enterprises on responsible AI adoption, governance frameworks, and enterprise transformation. Yet their own published research is now being scrutinized for exactly the types of AI failures they warn clients about.
What Happened?
According to an investigation conducted by GPTZero researchers and independently verified by the Financial Times, four reports produced by PwC Middle East over the past two years contained numerous citation and sourcing problems. The affected publications covered topics including agentic AI, digital government services, autonomous systems, and the future of electric and autonomous vehicles.
Rather than containing isolated typographical mistakes, researchers identified patterns commonly associated with AI-generated content that had not undergone comprehensive editorial review.
Among the reported issues were:
- References pointing to webpages that either did not exist or contained unrelated information.
- Citations supporting claims that could not be verified.
- An academic paper cited as evidence despite researchers finding no record that it had ever been published.
- Incorrect attribution of sources.
- Broken hyperlinks throughout multiple reports.
- References that linked to material which never supported the accompanying statement.
Perhaps the most unusual discovery involved one citation URL that still contained the parameter “utm_source=chatgpt.com”, strongly suggesting that content or references had been copied directly from AI-generated output without sufficient review before publication.
Researchers also noted that one report cited a teenage Medium blogger as supporting evidence for what PwC described as a real-world enterprise AI success story involving JPMorgan. While the underlying JPMorgan initiative itself is genuine, the cited source was an unusual choice for a consulting report intended for governments and enterprise executives.
PwC Responds
PwC Middle East acknowledged the issues after the investigation became public.

The firm stated that it takes the accuracy of its published research seriously and is updating a limited number of supporting citations. It also said it has established quality-control processes governing research and content development that employees are expected to follow. The company did not publicly explain how the problematic citations passed internal review before publication.
📬 Stay Ahead of Cyber Threats
Get the latest cybersecurity news, critical vulnerabilities, threat intelligence, tutorials, and exclusive giveaways delivered straight to your inbox. No spam. Unsubscribe anytime.
Subscribe to the Newsletter →At the time of writing, there has been no indication that the reports were entirely withdrawn. Instead, PwC has indicated that corrections are being made to affected citations.
Not the First Big Four AI Embarrassment
The incident follows earlier investigations into reports published by other Big Four consulting firms.
Earlier this year, GPTZero researchers identified similar citation problems in a KPMG report discussing agentic AI. The report reportedly contained fabricated references, inaccurate citations, and multiple hallucinated sources, ultimately leading KPMG to withdraw the publication.
EY also faced criticism after separate investigations uncovered AI-generated inaccuracies within published research, prompting revisions and renewed attention toward editorial oversight. With PwC now joining that list, all three incidents point toward a common pattern rather than isolated human error.
Why AI Produces Fake Citations
The technical problem behind these incidents is well understood within the machine learning community.

Large language models do not function as search engines or academic databases. They generate text by predicting statistically probable sequences of words based on patterns learned during training. Although modern models can accurately summarize information, they do not inherently verify whether a cited paper, author, DOI, or webpage actually exists unless they are explicitly connected to authoritative retrieval systems. This phenomenon is commonly referred to as AI hallucination.
When asked to produce research reports complete with references, an LLM may generate citation structures that appear perfectly legitimate. The title sounds plausible. The journal name exists. The author names resemble real researchers. The publication year fits the context. However, the cited paper itself may never have existed. This happens because the model is optimizing for linguistic probability rather than factual validation.
Citation hallucinations generally fall into several categories:
- Completely fabricated academic papers.
- Real papers paired with incorrect conclusions.
- Valid authors attached to nonexistent publications.
- Correct URLs that lead to unrelated pages.
- Broken or malformed hyperlinks.
- References copied from previous AI conversations without verification.
Unlike a traditional search engine, an LLM has no built-in obligation to reject uncertain information. Unless constrained by retrieval-augmented generation (RAG), external databases, or manual verification workflows, it may confidently produce convincing but incorrect citations.
Why the “utm_source=chatgpt.com” Parameter Matters
Among all the reported findings, the appearance of “utm_source=chatgpt.com” within one citation attracted particular attention. UTM parameters are marketing tags appended to URLs for analytics purposes. Their presence indicates where traffic originated.
Finding such a parameter inside a published citation strongly suggests that the URL may have been copied directly from AI-generated output rather than independently sourced through conventional editorial research. Although the parameter alone does not prove the entire report was AI-written, it provides unusually direct evidence that generative AI likely played a role in assembling at least part of the source material.
The Editorial Patterns Researchers Observed
Beyond fabricated references, GPTZero researchers reportedly identified structural patterns that frequently appear in AI-generated documents.
One example involved the same factual claim being repeated multiple times within a report while being supported by different citations each time. Researchers argued this inconsistency reflects statistical text generation rather than careful human sourcing. Another example involved an internal PwC survey referenced in the report, while the accompanying citation instead linked to an unrelated media article that contained none of the referenced survey data. Individually, these issues might resemble editorial oversight. Collectively, they point toward systemic failures in source verification.
Why This Matters Beyond PwC
Thought leadership reports occupy an unusual position within enterprise decision making.
Unlike marketing blogs, these publications frequently influence procurement decisions, digital transformation strategies, public policy discussions, and boardroom investments. Government agencies, regulators, investors, and corporate executives often reference them when evaluating emerging technologies.
When citations inside those reports cannot be verified, the consequences extend beyond reputational damage. Organizations may make strategic decisions based on unsupported evidence. Market forecasts become difficult to validate. Academic credibility is weakened. Public trust in AI-assisted research declines.
These risks become particularly significant when reports discuss rapidly evolving technologies such as autonomous AI agents, cybersecurity, electric vehicles, or national AI policy.
The Technical Solution Is Not Simply “Use Less AI”
The incident should not be interpreted as evidence that generative AI cannot assist research. Instead, it highlights the difference between content generation and knowledge verification.
Enterprise AI systems are increasingly moving toward architectures that combine language models with retrieval systems capable of querying verified knowledge bases in real time. Retrieval-Augmented Generation (RAG), citation validation pipelines, DOI verification, URL checking, automated link integrity testing, and human editorial review all reduce hallucination risk significantly.
Many organizations are also introducing multi-stage review pipelines where AI drafts material, automated tools verify references, and subject-matter experts perform final validation before publication. For consulting firms producing research intended to influence billion-dollar decisions, these controls are rapidly becoming essential rather than optional.
A Growing Governance Challenge
The repeated appearance of hallucinated citations across multiple consulting firms reflects a broader governance issue rather than a flaw unique to any single organization.
As firms race to integrate generative AI into internal workflows, productivity gains can easily outpace quality assurance processes. AI dramatically accelerates drafting, summarization, and report production, but it does not eliminate the responsibility to verify every factual claim, statistic, hyperlink, and citation before publication.
The more authoritative the publication appears, the higher the standard for verification becomes. That expectation is especially relevant for organizations advising clients on responsible AI deployment.
Conclusion
PwC’s citation problems may ultimately be corrected through revised reports, but the incident reinforces a lesson that continues to surface across industries: generative AI is an effective drafting assistant, not an autonomous research authority.
The discovery of fabricated references, nonexistent studies, broken links, inconsistent sourcing, and a citation containing “utm_source=chatgpt.com” demonstrates how easily convincing documents can conceal factual weaknesses when human verification is insufficient.
For organizations adopting AI at scale, the takeaway is straightforward. The challenge is no longer generating professional-looking reports. It is ensuring every citation, every statistic, and every conclusion can withstand independent verification. As more enterprises integrate AI into research workflows, rigorous editorial controls will increasingly determine whether AI enhances credibility or quietly undermines it.









