Google’s unreleased Gemini 4 lineup just leaked—and the internal benchmarks look wild.
According to internal chats and documents obtained by Business Insider, Google is testing a new model variant codenamed "Carbon" that reportedly rivals Anthropic’s Opus 5.5 in coding performance.
Here is everything you need to know:
2/7 The Element Codename Scheme
Google is actively testing several internal Gemini 4 branches:
Carbon was recently pushed to Google’s internal developer platform, Jetski.
Initial employee feedback shows Carbon significantly outpaces baseline Argon on programming tasks. One Googler explicitly compared its coding output to Claude Opus 5.5—a notable leap from earlier Argon checkpoints, which felt closer to Opus 5.
4/7 Where Does It Fit in the Gemini Lineup?
Google previously categorized its model tiers by role:
5/7 Driven by "Recursive Self-Improvement"?
Why are iterations shipping this fast before Gemini 4 is officially released?
Responding to the leak, Google DeepMind researcher Vedant Misra teased: "Have you heard of recursive self-improvement?"
Google, Anthropic, and OpenAI are increasingly using frontier reasoning models to train, evaluate, and optimize their next generations.
6/7 Gemini 4 Launch Prep Spotted in the Wild
Google is already rewiring its public-facing interfaces for the rollout:
The race for frontier coding dominance is accelerating. If Carbon delivers Opus 5.5-level coding at Google's scale and native integration, the upcoming Gemini 4 launch could shift the leaderboard once again.
Source: You do not have permission to view the full content of this post. Log in or register now.
According to internal chats and documents obtained by Business Insider, Google is testing a new model variant codenamed "Carbon" that reportedly rivals Anthropic’s Opus 5.5 in coding performance.
Here is everything you need to know:
2/7 The Element Codename Scheme
Google is actively testing several internal Gemini 4 branches:
- Argon: Google’s flagship frontier reasoning model.
- Barium: An earlier checkpoint (Argon was previously called Barium-B).
- Carbon: The newest internal build focused on elite software engineering capabilities.
Carbon was recently pushed to Google’s internal developer platform, Jetski.
Initial employee feedback shows Carbon significantly outpaces baseline Argon on programming tasks. One Googler explicitly compared its coding output to Claude Opus 5.5—a notable leap from earlier Argon checkpoints, which felt closer to Opus 5.
4/7 Where Does It Fit in the Gemini Lineup?
Google previously categorized its model tiers by role:
- Argon: Frontier reasoning and complex logic
- Flash: Low latency and high throughput
- Omni: Multimodal generative media
- Gemma: Open-weights edge execution
5/7 Driven by "Recursive Self-Improvement"?
Why are iterations shipping this fast before Gemini 4 is officially released?
Responding to the leak, Google DeepMind researcher Vedant Misra teased: "Have you heard of recursive self-improvement?"
Google, Anthropic, and OpenAI are increasingly using frontier reasoning models to train, evaluate, and optimize their next generations.
6/7 Gemini 4 Launch Prep Spotted in the Wild
Google is already rewiring its public-facing interfaces for the rollout:
- Gemini App: New "Automatic" model routing and dynamic reasoning sliders (Low to High).
- AI Studio: A spotted "Ultra" mode offering advanced tools and skills.
- Google’s Logan Kilpatrick confirmed the team is actively optimizing to squeeze maximum performance out of Argon.
The race for frontier coding dominance is accelerating. If Carbon delivers Opus 5.5-level coding at Google's scale and native integration, the upcoming Gemini 4 launch could shift the leaderboard once again.
Source: You do not have permission to view the full content of this post. Log in or register now.