Did OpenAI’s Group Theory Breakthrough Come From a Mathematician’s Private ChatGPT Conversations?

The CyberSec Guru

Did OpenAI Use a Mathematician’s Private ChatGPT Chats

If you like this post, then please share it:

Buy me A Coffee!

Support The CyberSec Guru’s Mission

🔐 Fuel the cybersecurity crusade by buying me a coffee! Why your support matters: Zero paywalls: Keep the main content 100% free for learners worldwide.

“Your coffee keeps the servers running and the knowledge flowing in our fight against cybercrime.”☕ Support My Work

Buy Me a Coffee Button

TU Dresden professor Andreas Thom has published an email exchange alleging that OpenAI’s categorical denial (“that did not happen”) erased exactly the distinction his question drew, and that the company’s own language in a separate dispute contradicts it. Here is the full technical picture: the mathematics at stake, the three very different mechanisms his email separated, and what evidence would actually settle the question.

Key takeaways

  • OpenAI announced last month that its Astra model settled a 27-year-old open problem in group theory, the existence of a non-sofic group, with the proof’s decisive step building on methods Andreas Thom published with Gábor Kun.
  • In the months before that announcement, Thom and a colleague in Dresden had been actively working through the same expander-matching mathematics with ChatGPT. He emailed OpenAI researchers Mark Sellke and Sébastien Bubeck asking whether those conversations entered training data, and whether they were accessible to the model while it worked.
  • Sellke’s complete reply, as Thom has now published it: “Regarding your conversations with ChatGPT: that did not happen.” No qualification, no explanation, no evidence, no account-specific check.
  • Weeks later, in the separate Buckmaster–Alpöge Navier–Stokes dispute, OpenAI stated that no specific user data was accessed but that it “cannot rule out that de-identified data derived from their usage of our products helped improve our models.” That is precisely the caveat absent from the answer Thom received.
  • The dispute hinges on a technical distinction most coverage misses: verbatim ingestion into training data, inference-time access to conversation history, and distilled capability from de-identified usage data are three different mechanisms with three different audit trails. Only OpenAI holds the logs to distinguish them.

The announcement that started it

Last month OpenAI claimed that Astra, its flagship reasoning model, had settled a question that had been open for roughly 27 years: whether there exists a group that is not sofic. Sofic groups sit at a crossroads of geometric group theory, operator algebras and dynamical systems, and the question of whether all countable groups are sofic has resisted every attack since Misha Gromov and Alan Weiss put the framework on the map around the turn of the millennium. A positive answer would have been surprising; a construction of a counterexample is the result most specialists spent careers circling.

The reason the announcement landed with particular force in one Dresden office is that the proof’s key step, by OpenAI’s own framing, came out of the expander-based program Andreas Thom published with Gábor Kun of the Rényi Institute in Budapest. Thom, a professor at Technische Universität Dresden with a long record in soficity and related approximation problems, was not an idle bystander to that program. According to his own account, he and a Dresden colleague had spent the preceding months actively working through the expander matching problem and various extensions of the Kun–Thom work, using ChatGPT as a working partner to probe definitions, test obstructions, and draft arguments.

So when the announcement appeared, Thom’s question was not primarily “is the proof correct?” (verification of a claimed non-soficity construction is a slow, communal process still underway) but “where did the model’s route come from?” He emailed Mark Sellke, the probabilist-turned-OpenAI-researcher credited on the work, and Sébastien Bubeck, who has led much of OpenAI’s mathematical reasoning effort, and he has now published both the email and the reply.

The email exchange: two questions, one five-word answer

Thom’s email, quoted in his Mastodon thread, is careful. He writes that he and a colleague in Dresden “were discussing the expander matching problem and various extensions of the work with Gabor Kun actively over the last months with ChatGPT, so that we are of course curious if that was part of the training data or accessible to the reasoning process.” He adds that “there is a certain (frankly unacceptable) lack of transparency here; and I fear it will damage the communal process of math more than the new AI-generated results will benefit the subject.”

Parse that sentence and you see two structurally different questions. Question one: did the content of those conversations enter the datasets used to train subsequent models? Question two: were those conversations accessible to the model at inference time, through account memory, retrieval over chat history, or any other mechanism, while it was working on the non-soficity problem? The first is a question about dataset lineage; the second is a question about inference logs. They have different answers, different audit trails, and different ethical implications.

Mark Sellke’s complete answer, as Thom reports it, was a single sentence: “Regarding your conversations with ChatGPT: that did not happen.”

Thom’s reading, set out across a three-part thread this week, is that the categorical denial can only coherently address the second question, direct access during solving, because the first question cannot be answered categorically without inspecting training pipelines that only OpenAI can inspect. “The categorical answer now looks as though it addressed only direct access under (2). No such qualification, explanation, or evidence was given. I take this as dishonesty to say the least.” In his third post he sharpens it: the answer was “at minimum, unjustifiably broad and materially misleading; looking back it was plainly dishonest.”

Sofic groups, explained without hand-waving

A countable group is sofic if, roughly, every finite piece of its multiplication table can be modeled by permutations of a finite set, up to arbitrarily small error. Formally: for every finite subset F of the group and every ε > 0, there is a map from F into the symmetric group Sym(n) for some n, such that the map is multiplicative on F up to ε in the normalized Hamming distance, and distinct elements of F have images that disagree on almost all points. An intuitive picture: a sofic group is one whose every finite “shadow” can be faked by a finite permutation system. Amenable groups are sofic, residually amenable groups are sofic, linear groups are sofic, and the class is closed under many natural operations, which is exactly why finding a counterexample is hard. Every candidate construction must defeat not one approximation but all possible approximations simultaneously.

📬 Stay Ahead of Cyber Threats

Get the latest cybersecurity news, critical vulnerabilities, threat intelligence, tutorials, and exclusive giveaways delivered straight to your inbox. No spam. Unsubscribe anytime.

Subscribe to the Newsletter →

The stakes are not merely taxonomic. Soficity implies several famous conjectures, including Kaplansky’s direct finiteness conjecture for group rings and Gottschalk’s surjunctivity conjecture; a genuine non-sofic group would settle those lines of questioning at a stroke and reshape how group theorists think about finite approximability. That is why a claimed construction is a headline, and why provenance matters: a result this consequential carries careers, credit, and priority with it.

The two research programs competing to break soficity

Since around 2020, the field’s energy has split between two very different attack vectors. The combinatorial route, associated with Kun and Thom, engineers group relations so that any putative sofic approximation would manufacture a family of bounded-degree expander graphs carrying a matching or decomposition structure that expanders provably resist. Expansion, measured by the spectral gap, forbids certain balanced combinatorial configurations, and the program tries to force a group whose finite models would require exactly those configurations. It is elegant and it is a long shot; the obstruction has to survive every possible way of cheating with ultraproducts and non-uniform approximations.

The second route runs through quantum information and operator algebras: non-local games, commuting-operator versus tensor-product correlations, Thomas Slofstra’s solution groups, and the shockwave of the MIP* = RE theorem of Ji, Natarajan, Vidick, Wright and Yuen, which refuted the Connes embedding conjecture and demonstrated that obstruction machinery of exactly the right flavor exists in quantum correlation theory. By Thom’s own assessment, this was the main line of attack: “The approach of Kun and myself was not the main line of attack on non-soficity, in fact there were other more promising approaches along the line of quantum games etc.”

That assessment is what makes the provenance question sharp. If a model produces a proof displaying detailed command of the less-traveled route, including moves that live in unpublished extensions of it worked through in private chat sessions, then the surprising fact is not merely that a proof exists. It is that an autonomous search landed precisely where one specific, private research program had been walking, rather than on the route the visible literature rated as more promising. Thom puts it plainly: “OpenAI’s detailed command of our techniques therefore made me wonder how the model found this route.”

Training data, inference-time access, and distilled capability: three different questions

Most commentary on this dispute conflates mechanisms that machine-learning practitioners keep strictly separate, and the conflation is doing real damage to the public discussion. There are at least three ways private conversation content could influence a model.

The first is verbatim or near-verbatim ingestion into training data: retained logs are curated, filtered, de-identified or not, and mixed into pretraining or post-training corpora. This is the mechanism the memorization literature makes concrete. Carlini and colleagues’ 2021 extraction attacks showed that large language models can regurgitate training substrings verbatim, and membership-inference techniques can sometimes detect whether a given document was in a mixture. Answering Thom’s question one definitively requires dataset cards, data lineage, and checkpoint-level diffs: which datasets fed which checkpoint, and whether his sessions appear in any of them.

The second is inference-time access: account-level memory features, retrieval augmented over a user’s own history, or context injected by tooling. This is a question about serving logs and retrieval indices, and it is the question a denial like “that did not happen” most naturally reads as addressing: the model did not have Thom’s chats in context while solving.

The third is the subtle one, and it is the one OpenAI’s own later language preserves: distilled capability. Aggregated, de-identified usage data (preference signals, interaction traces, curated session derivatives) can shape post-training behavior without any verbatim memorization and without any inference-time lookup. An idea can enter a model’s policy as a tendency, a technique, a way of framing an expander obstruction, without any string from the original conversation surviving. This mechanism is neither “training data” in the naive verbatim sense nor “access” in the inference sense, and it is exactly what OpenAI conceded it “cannot rule out” in the Buckmaster–Alpöge case.

This is why Thom’s charge of overbreadth is technically coherent rather than merely aggrieved. A categorical denial spanning all three mechanisms, issued without disclosing product and privacy settings in effect, relevant datasets and checkpoints, or what “de-identified data derived from usage” operationally means, asserts knowledge that no external party can verify and that the issuer can only possess through an internal lineage audit it did not describe. As Thom writes: “We are not required to reverse-engineer OpenAI’s internal training pipeline to establish what happened. Only OpenAI has the relevant data for that. If Sellke and Bubeck give a categorical denial, it must disclose the basis.”

The opt-out question compounds it. Thom disabled model training on his conversations on 29 June. That control, the ChatGPT training toggle familiar to every user, is prospective: it governs future use of future conversations. It says nothing about sessions already retained, already curated into a dataset, or already distilled into a checkpoint, and its implementation is a promise whose enforcement users cannot audit. “That control is still only a promise whose implementation users cannot audit, and it is prospective: it does not answer what happened to earlier conversations or to derivatives already selected. OpenAI’s answer did not mention the setting or any account-specific check.”

And de-identification, in this specific domain, carries a peculiarity worth stating clearly: mathematical ideas are themselves identifying. Stripping a name and an account ID from a discussion of a novel expander-matching obstruction does not anonymize the discussion, because the idea is unique enough to fingerprint its originators to anyone in the field. “De-identification may remove a name; it does not remove the intellectual content of a mathematical idea.”

The Navier–Stokes parallel: Buckmaster–Alpöge and the missing caveat

Weeks after Thom received his five-word answer, OpenAI found itself in an adjacent dispute over claimed progress on the Navier–Stokes regularity problem, involving mathematicians Tristan Buckmaster and Lena Alpöge. There, the company’s public position was notably more careful: no specific user data was accessed, but it “cannot rule out that de-identified data derived from their usage of our products helped improve our models.”

Set the two responses side by side and the asymmetry is the story. “OpenAI is drawing a distinction that its answer to me erased, despite the fact that my question explicitly made that distinction,” Thom writes. Either the epistemic situation in the June-to-announcement window was cleaner than in the Navier–Stokes window (in which case the basis for that cleanliness is exactly what was never disclosed), or the same irreducible uncertainty applied to both, and only one correspondent was told about it.

Thom draws a further inference from conduct around the proposed papers themselves. In his third post he cites “Bubeck’s acknowledged attacks and his attempt to exclude Alpöge from a proposed OpenAI-proof paper” as deepening “the concern that there is a loss of moral compass. Taken together, this raises serious questions about their judgment and personal integrity.” In mathematics, the boundary between authorship, acknowledgment and silence is the discipline’s moral economy: an “acknowledged attack” borrowed into a proof is credit owed, and excluding a contributor from a paper that uses their lines of attack is not a clerical choice. Whether or not any individual allegation is ultimately substantiated, the pattern Thom describes (categorical denial to one correspondent, hedged admission to another, contested credit in between) is what makes this more than a single email gone wrong.

What would real evidence look like?

It is worth being precise about what would settle this, because “trust me” in either direction is not evidence. Thom’s list is a reasonable starting point and maps cleanly onto standard ML governance artifacts: the product and privacy settings in effect on his account across the relevant window; the datasets and checkpoints used for any model touched by the non-soficity work, with lineage sufficient to check membership; an operational definition of “de-identified data derived from usage”; and an account-specific check of retrieval, memory and serving logs for the solving runs. None of this is exotic. Dataset cards, model cards and system cards exist precisely to make lineage claims inspectable, and litigation discovery in copyright suits against AI vendors has already demonstrated that retention schedules, log pipelines and deletion events are trackable internal facts, not mysteries.

Two cautions cut in opposite directions, and a serious article should state both. Absence of disclosure is not proof of misuse: the Kun–Thom program is published, any competent team could have built on it legitimately, and Sellke and Bubeck are sophisticated mathematicians who did not need a chat log to know where the expander route leads. A model reaching a minority route through massive search and reinforcement learning on proof attempts would be surprising but not impossible. Equally, a categorical denial issued without basis is not proof of cleanliness: the honest epistemic state, given how modern training pipelines consume derived data, is the hedged one OpenAI itself used elsewhere. The failure here is not necessarily that something was taken. It is that the confidence of the answer exceeded the evidence anyone was shown.

Why mathematicians should care even if nothing leaked

Strip the personalities out and the structural problem remains. Research assistants built on LLMs are genuinely useful; many working mathematicians use them daily for literature triage, example generation and error-checking, and nothing in Thom’s thread argues for abstinence. What his thread argues for is governance of the boundary between private and public research. If nonpublic ideas supplied by users can improve a model, and the provider can then deploy that model to race those same users to publication (without informed consent, without disclosure, without credit), the communal process that makes mathematics self-correcting has been privatized on one side. “I fear it will damage the communal process of math more than the new AI-generated results will benefit the subject,” Thom wrote before he ever received a reply, and the subsequent exchange has only widened the audience for that fear. The predictable downstream effect is a chilling one: researchers self-censoring their best half-formed ideas out of tools that could otherwise accelerate them, which helps nobody, least of all the science.

OpenAI’s position, and what it has not said

As of publication, the public artifacts on OpenAI’s side are three: Sellke’s one-sentence email, the hedged statement in the Buckmaster–Alpöge Navier–Stokes dispute, and the company’s standing privacy documentation describing de-identified usage data and user training controls. There is no published account-specific check of Thom’s account, no dataset or checkpoint lineage addressing his questions, and no public response to his thread. That silence is itself a data point in a dispute whose entire subject is transparency.

FAQ

Did OpenAI admit training on Andreas Thom’s ChatGPT conversations? No. OpenAI’s direct answer to Thom was a categorical denial (“that did not happen”), and its later public statement in a separate dispute said no specific user data was accessed while conceding it cannot rule out that de-identified usage-derived data improved its models. Thom’s allegation is that the denial was overbroad and misleading, not that OpenAI confessed.

What is a sofic group in plain English? A group whose every finite fragment of its multiplication structure can be approximated by permutations of a finite set, with arbitrarily small error. Think of it as a group all of whose finite “photographs” can be convincingly staged by finite permutation systems; whether every countable group has this property has been open for roughly 27 years.

What exactly were Thom’s two questions? First, whether his ChatGPT conversations entered training data; second, whether they were accessible to the model during its reasoning on the non-soficity problem. The first is a dataset-lineage question, the second an inference-logs question; they require different evidence to answer.

What does “de-identified data derived from usage” mean? Aggregated or stripped-down derivatives of user interactions, traces, preferences, curated session data, used to improve models without attached identifiers. Thom’s point is that in mathematics, removing names does not remove the identifying intellectual content of an idea.

Does disabling ChatGPT training protect past conversations? No. The training toggle is prospective: it governs future conversations. It does not reach back to sessions already retained, already included in datasets, or already distilled into model weights, and users cannot audit its implementation.

Has the non-sofic group proof been verified? Independent verification of a claimed construction of this difficulty is a slow communal process, and the provenance dispute documented here is separate from, though entangled with, the mathematical review now underway.

Bottom line

This controversy will be remembered less for one group-theoretic claim than for what it exposes about auditability in the LLM era: the boundary between a researcher’s private thinking and a vendor’s training pipeline is currently enforced by promises, prospective toggles, and five-word emails. Andreas Thom’s published exchange does not prove that his conversations were used; what it proves is that when a mathematician asked the two questions that matter, the company with exclusive access to the answers returned a sentence whose confidence outran its evidence, and then, weeks later, in a different dispute with different reputational stakes, supplied the very caveat it had withheld from him. In a discipline where precisely calibrated confidence is the whole point, that asymmetry is the result worth publishing.

Sources and primary documents: Andreas Thom’s three-part Mastodon thread (@andreasthom@mathstodon.xyz, replying to @tristanbuckmaster) and the email exchange published therein; OpenAI’s public statement in the Buckmaster–Alpöge Navier–Stokes dispute; the Kun–Thom published work on expander-based approaches to soficity; Ji, Natarajan, Vidick, Wright and Yuen, “MIP* = RE”; Slofstra’s work on solution groups of non-local games; Carlini et al., “Extracting Training Data from Large Language Models” (2021). Allegations attributed to individuals are reported as stated by those individuals and, where applicable, as denied or qualified by OpenAI.

Buy me A Coffee!

Support The CyberSec Guru’s Mission

🔐 Fuel the cybersecurity crusade by buying me a coffee! Your contribution powers free tutorials, hands-on labs, and security resources.

Why your support matters:
  • Writeup Access: Get complete writeup access within 12 hours
  • Zero paywalls: Keep the main content 100% free for learners worldwide

Perks for one-time supporters:
☕️ $5: Shoutout in Buy Me a Coffee
🛡️ $8: Fast-track Access to Live Webinars
💻 $10: Vote on future tutorial topics + exclusive AMA access

“Your coffee keeps the servers running and the knowledge flowing in our fight against cybercrime.”☕ Support My Work

Buy Me a Coffee Button

If you like this post, then please share it:

Analysis

Discover more from The CyberSec Guru

Subscribe to get the latest posts sent to your email!

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from The CyberSec Guru

Subscribe now to keep reading and get access to the full archive.

Continue reading