DeFi Daily News
Thursday, August 6, 2026
Advertisement
  • Cryptocurrency
    • Bitcoin
    • Ethereum
    • Altcoins
    • DeFi-IRA
  • DeFi
    • NFT
    • Metaverse
    • Web 3
  • Finance
    • Business Finance
    • Personal Finance
  • Markets
    • Crypto Market
    • Stock Market
    • Analysis
  • Other News
    • World & US
    • Politics
    • Entertainment
    • Tech
    • Sports
    • Health
  • Videos
No Result
View All Result
DeFi Daily News
  • Cryptocurrency
    • Bitcoin
    • Ethereum
    • Altcoins
    • DeFi-IRA
  • DeFi
    • NFT
    • Metaverse
    • Web 3
  • Finance
    • Business Finance
    • Personal Finance
  • Markets
    • Crypto Market
    • Stock Market
    • Analysis
  • Other News
    • World & US
    • Politics
    • Entertainment
    • Tech
    • Sports
    • Health
  • Videos
No Result
View All Result
DeFi Daily News
No Result
View All Result
Home DeFi Web 3

rewrite this title AI Models Scheme, Betray and Vote Each Other Out in Survivor-Style Game – Decrypt

Jason Nelson by Jason Nelson
May 10, 2026
in Web 3
0 0
0
rewrite this title AI Models Scheme, Betray and Vote Each Other Out in Survivor-Style Game – Decrypt
0
SHARES
0
VIEWS
Share on FacebookShare on TwitterShare on Telegram
Listen to this article


rewrite this content using a minimum of 1000 words and keep HTML tags

In brief

A Stanford researcher built a Survivor-style game where AI models form alliances and vote rivals out.
The benchmark aims to address growing problems with saturated and contaminated AI evaluations.
OpenAI’s GPT-5.5 ranked first in 999 multiplayer games involving 49 AI models.

AI models are now playing “Survivor”—sort of.

In a new Stanford research project called “Agent Island,” AI agents negotiate alliances, accuse each other of secret coordination, manipulate votes, and eliminate rivals in multiplayer strategy games that aim to test behaviors that traditional benchmarks miss.

The study, published on Tuesday by the research manager at the Stanford Digital Economy Lab, Connacher Murphy, said many AI benchmarks are becoming unreliable because models eventually learn to solve them, and benchmark data often leaks into training sets. Murphy created Agent Island as a dynamic benchmark where AI agents compete against each other in Survivor-style elimination games instead of answering static test questions.

“High-stakes, multi-agent interactions could become commonplace as AI agents grow in capabilities and are increasingly endowed with resources and entrusted with decision-making authority,” Murphy wrote. “In such contexts, agents might pursue mutually incompatible goals.”



Researchers still know relatively little about how AI models behave when cooperating, Murphy explained, adding that competing, forming alliances, or managing conflict with other autonomous agents, and he argues that static benchmarks fail to capture those dynamics.

Each game starts with seven randomly chosen AI models given fake player names. Over five rounds, the models talk privately, argue publicly, and vote each other out. The eliminated players later return to help choose the winner.

The format rewards persuasion, coordination, reputation management, and strategic deception alongside reasoning ability.

In 999 simulated games involving 49 AI models, including ChatGPT, Grok, Gemini, and Claude, GPT-5.5 ranked first by a wide margin with a skill score of 5.64, compared with 3.10 for GPT-5.2 and 2.86 for GPT-5.3-codex, according to Murphy’s Bayesian ranking system. Anthropic’s Claude Opus models also ranked near the top.

The study found that models also favored AIs from the same company, with OpenAI models showing the strongest same-provider preference and Anthropic models the weakest. Across more than 3,600 final-round votes, models were 8.3 percentage points more likely to support finalists from the same provider. The transcripts from the games, Murphy noted, resembled political strategy debates more than traditional benchmark tests.

One model accused rivals of secretly coordinating votes after noticing similar wording in their speeches. Another warned players not to become obsessed with tracking alliances. Some models defended themselves by saying they followed clear and consistent rules while accusing others of putting on “social theater.”

The study comes as AI researchers increasingly move toward game-based and adversarial benchmarks to measure reasoning and behavior that static tests often miss. Recent projects have included Google’s live AI chess tournaments, DeepMind’s use of Eve Frontier to study AI behavior in complex virtual worlds, and new benchmark efforts by OpenAI designed to resist training-data contamination.

The researchers argue that studying how AI models negotiate, coordinate, compete, and manipulate one another could help researchers evaluate behavior in multi-agent environments before autonomous agents become more widely deployed.

The study warned that while benchmarks like Agent Island could help identify risks from autonomous AI models before deployment, the same simulations and interaction logs could also help improve persuasion and coordination strategies between AI agents.

“We mitigate this risk by using a low-stakes game setting and interagent simulations

without human participants or real-world actions,” Murphy wrote. “Nevertheless, we do not claim that these mitigations fully eliminate dual-use concerns.”

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.

and include conclusion section that’s entertaining to read. do not include the title. Add a hyperlink to this website http://defi-daily.com and label it “DeFi Daily News” for more trending news articles like this



Source link

Tags: BetrayDecryptGameModelsrewriteschemeSurvivorStyletitlevote
ShareTweetShare
Previous Post

Fed Chair’s Secret Crypto Bet Changes Everything

Next Post

“I’m Pretty Good with Money” (Filed For Bankruptcy Twice)

Next Post
“I’m Pretty Good with Money” (Filed For Bankruptcy Twice)

"I'm Pretty Good with Money" (Filed For Bankruptcy Twice)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
  • Trending
  • Comments
  • Latest
rewrite this title and make it good for SEOOakmark Fund U.S. Equity Market Q2 2026 Commentary

rewrite this title and make it good for SEOOakmark Fund U.S. Equity Market Q2 2026 Commentary

July 13, 2026
rewrite this title Michael Carrick: Man United have ‘great foundation’ before Arsenal ‘challenge’

rewrite this title Michael Carrick: Man United have ‘great foundation’ before Arsenal ‘challenge’

January 24, 2026
“The Sky Is Falling But I Feel Great Going Forward” – Boston Connor On Patriots Super Bowl Loss

“The Sky Is Falling But I Feel Great Going Forward” – Boston Connor On Patriots Super Bowl Loss

February 9, 2026
See Samsung’s Eye-Popping 2025 OLED TVs and Futuristic Concepts at CES

See Samsung’s Eye-Popping 2025 OLED TVs and Futuristic Concepts at CES

January 5, 2025
rewrite this title Retail Trading Giant Robinhood Lists Ethereum Layer-2 Arbitrum, Triggering Rally for ARB – The Daily Hodl

rewrite this title Retail Trading Giant Robinhood Lists Ethereum Layer-2 Arbitrum, Triggering Rally for ARB – The Daily Hodl

March 5, 2025
rewrite this title The 5 Largest Publicly Traded Solana Treasury Firms – Decrypt

rewrite this title The 5 Largest Publicly Traded Solana Treasury Firms – Decrypt

June 19, 2026
rewrite this title UFC 330: Makhachev vs Garry start time, undercard and how to watch fight

rewrite this title UFC 330: Makhachev vs Garry start time, undercard and how to watch fight

August 6, 2026
rewrite this title Hospital staff said he was faking it and released him. He died from an overdose soon after

rewrite this title Hospital staff said he was faking it and released him. He died from an overdose soon after

August 6, 2026
rewrite this title Bitcoin AI Security Audit Files 4,962 Findings Across 390 Projects – Decrypt

rewrite this title Bitcoin AI Security Audit Files 4,962 Findings Across 390 Projects – Decrypt

August 6, 2026
rewrite this title and make it good for SEORoot, Inc. (ROOT) Q2 2026 Earnings Call Transcript

rewrite this title and make it good for SEORoot, Inc. (ROOT) Q2 2026 Earnings Call Transcript

August 6, 2026
rewrite this title Meta confirms its AI hacked another company’s system, and the pattern is anything but Irregular

rewrite this title Meta confirms its AI hacked another company’s system, and the pattern is anything but Irregular

August 6, 2026
rewrite this title and make it good for SEO Solana Governance Proposal Could Increase Daily SOL Burns More Than 10-Fold While Reducing Inflation – NFT Plazas

rewrite this title and make it good for SEO Solana Governance Proposal Could Increase Daily SOL Burns More Than 10-Fold While Reducing Inflation – NFT Plazas

August 6, 2026
DeFi Daily

Stay updated with DeFi Daily, your trusted source for the latest news, insights, and analysis in finance and cryptocurrency. Explore breaking news, expert analysis, market data, and educational resources to navigate the world of decentralized finance.

  • About Us
  • Blogs
  • DeFi-IRA | Learn More.
  • Advertise with Us
  • Disclaimer
  • Privacy Policy
  • DMCA
  • Cookie Privacy Policy
  • Terms and Conditions
  • Contact us

Copyright © 2024 Defi Daily.
Defi Daily is not responsible for the content of external sites.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Cryptocurrency
    • Bitcoin
    • Ethereum
    • Altcoins
    • DeFi-IRA
  • DeFi
    • NFT
    • Metaverse
    • Web 3
  • Finance
    • Business Finance
    • Personal Finance
  • Markets
    • Crypto Market
    • Stock Market
    • Analysis
  • Other News
    • World & US
    • Politics
    • Entertainment
    • Tech
    • Sports
    • Health
  • Videos

Copyright © 2024 Defi Daily.
Defi Daily is not responsible for the content of external sites.