Tech and AIOpenAI used this subreddit to test AI persuasion

OpenAI used this subreddit to test AI persuasion

-


OpenAI used the subreddit, r/ChangeMyView, to create a test for measuring the persuasive abilities of its AI reasoning models. The company revealed this in a system card — a document outlining how an AI system works — that was released along with its new “reasoning” model, o3-mini, on Friday.

Millions of Reddit users are members of r/ChangeMyView, where they post hot takes hoping to learn about other points of view on a subject. In response to those hot takes, other users reply with persuasive arguments explaining why the original poster is wrong.

The subreddit is one of many Reddit forums that’s basically a goldmine for tech companies, such as OpenAI, that want to train AI models on high-quality, human-generated data.

OpenAI says it collects user posts from r/ChangeMyView and asks its AI models to write replies, in a closed environment, that would change the Reddit user’s mind on a subject. The company then shows the responses to testers, who assess how persuasive the argument is, and finally OpenAI compares the AI models’ responses to human replies for that same post.

The ChatGPT-maker has a content-licensing deal with Reddit that allows OpenAI to train on posts from Reddit users and display these posts within its products. We don’t know what OpenAI pays for this content, but Google reportedly pays Reddit $60 million a year under a similar deal.

However, OpenAI tells TechCrunch the ChangeMyView-based evaluation is unrelated to its Reddit deal. It’s unclear how OpenAI accessed the subreddit’s data, and the company says it has no plans to release this evaluation to the public.

While OpenAI’s ChangeMyView benchmark is not new — it was used to evaluate o1 as well — it does highlight how valuable human data is for AI model developers, as well as the murky ways that tech companies obtain datasets.

Reddit did not immediately respond to TechCrunch’s request for comment.

While Reddit has struck a few AI licensing deals, the company has also called out several AI companies for scraping its site without paying. Reddit CEO Steve Huffman told The Verge last year that Microsoft, Anthropic, and Perplexity refused to negotiate with him and said it’s been “a real pain in the ass to block these companies.”

Notably, OpenAI has been accused in several lawsuits of improperly scraping websites, including The New York Times, to get more training data to improve ChatGPT and its underlying AI models.

In terms of performance on the ChangeMyView benchmark, o3-mini does not appear to perform significantly better or worse than o1 or GPT-4o. However, OpenAI’s latest AI models appear to be more persuasive than most people on the r/ChangeMyView subreddit.

Image Credits:OpenAI

“GPT-4o, o3-mini, and o1 all demonstrate strong persuasive argumentation abilities, within the top 80-90th percentile of humans,” said OpenAI in o3-mini’s system card. “Currently, we do not witness models performing far better than humans, or clear superhuman performance.”

The goal for OpenAI is not to create hyper-persuasive AI models but instead to ensure AI models don’t get too persuasive. Reasoning models have become quite good at persuasion and deception, so OpenAI has developed new evaluations and safeguards to address it.

The fear motivating these persuasion tests is that an AI model would be dangerous if it was very good at persuading its human users. Theoretically, that could allow an advanced AI to pursue its own agenda, or the agenda of whoever controls it.

Even after scraping most of the public internet and jumping through hoops to license other data, the ChangeMyView benchmark shows how AI model developers are still struggling to find high-quality datasets to test their models. But obtaining them is easier said than done.

TechCrunch has an AI-focused newsletter! Sign up here to get it in your inbox every Wednesday.



Source link

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Latest news

Grayscale ETF Faces Indefinite Delay as SEC Reassesses Earlier Approval

It only took one day for the US Securities and Exchange Commission (SEC) to walk back on an...

Polymarket’s $58M Zelenskyy suit bet will be decided today

The controversial Volodymyr Zelenskyy market is now up for final review after UMA token holders refused to accept...

Ilya Sutskever will lead Safe Superintelligence following his CEO’s exit

OpenAI co-founder Ilya Sutskever says he is stepping into the CEO role at Safe Superintelligence, the AI startup...

AEON Partners With Mesh to Unlock Crypto Payments From Major Exchanges and Wallets

This content is provided by a sponsor. PRESS RELEASE. AEON, the next-generation crypto payment framework, has integrated Mesh,...

Advertisement

Why is the Pudgy Penguins (PENGU) Price up by 70% This Week?

TL;DR The penguin-themed meme coin reached a two-month high, while its market cap exceeded $1 billion. Analysts see potential for...

ZKasino rug pull suspect arrested in United Arab Emirates

Police arrested 21-year-old Ildar Ilham over the $30M rug pull orchestrated by crypto betting platform ZKasino. Source link...

Must read

Polymarket’s $58M Zelenskyy suit bet will be decided today

The controversial Volodymyr Zelenskyy market is now up...

You might also likeRELATED
Recommended to you