Mathematicians want proof OpenAI didn’t use their work

Another researcher is challenging OpenAI about the data driving its increasingly impressive array of mathematical discoveries. Just days after a bitter row erupted over whether the company’s models benefited from unpublished work, a second mathematician has come forward accusing the AI giant of unethical and “dishonest” behavior and a lack of transparency about the origins of its training data.

In a series of posts on Mastodon, mathematician Andreas Thom raised concerns that interactions he and his colleagues had had with the ChatGPT chatbot before OpenAI’s triumphant announcement may have contributed to its success in the field. One of the 10 results OpenAI announced with great fanfare last month involved Thom’s area of expertise, so-called non-sofic groups, and OpenAI acknowledged that their result built heavily on previous work by Thom and fellow mathematician Gábor Kun.

Thom said he began reflecting on his own interactions with OpenAI after Tristan Buckmaster, a mathematics professor at New York University, publicly questioned whether the company’s AI models had benefited from his use of OpenAI’s Codex. After OpenAI announced its non-sofic groups result, it was widely criticized in mathematical circles for failing to acknowledge recent contributions from Thom and Kun and the company quietly amended its writeup. Non-sofic groups are, roughly speaking, infinite mathematical structures that cannot be approximated by finite ones.

Thom said he was also struck by “OpenAI’s detailed command of our techniques,” which he said were neither the most obvious nor the most promising routes to a solution at the time. He said he wrote emails to OpenAI researchers Sébastien Bubeck and Mark Sellke, also a statistician at Harvard, to ask whether his interactions with ChatGPT were “part of the training data or accessible to the reasoning process” and could therefore have contributed to the result.

But the answer did not satisfy Thom, who said it only addressed whether his conversations with the chatbot could be accessed directly, not whether they had entered into the vast pools of training data the company uses to improve its models. “No such qualification, explanation, or evidence was given,” he wrote. “I take this as dishonesty to say the least.”

Thom said researchers aren’t equipped to reverse-engineer OpenAI’s training pipeline to figure out whether their work has been used or not. “Only OpenAI has the relevant data for that.” If the company is going to deny doing this, he said the responsibility is on them to prove that by disclosing all necessary datasets and clarifying various settings and terms setting out how it uses data.

Keep reading

Unknown's avatar

Author: HP McLovincraft

Seeker of rabbit holes. Pessimist. Libertine. Contrarian. Your huckleberry. Possibly true tales of sanity-blasting horror also known as abject reality. Prepare yourself. Veteran of a thousand psychic wars. I have seen the fnords. Deplatformed on Tumblr and Twitter.

Leave a comment