Gemini 4 Pro Just Appeared in Disguise, and Its Early Tests Are Wild
Share this post:
Google may have finally shown its hand.
Over the past few days, a mysterious model labeled gemini-3.8-flash quietly appeared in Arena. That name should not have caused much excitement: the real Gemini 3.8 Flash had already launched earlier in September. But once developers began testing the new Arena variant, it became obvious that something was different.
The model was faster, more visually capable and far more comfortable with complex interactive work than people expected from a Flash release. After a wave of SVG, 3D, UI and game-generation tests, one theory quickly took over the AI community: this was not Gemini 3.8 Flash at all. It was Gemini 4 Pro wearing a temporary name.

The “gemini-3.8-flash” Model That Did Not Behave Like Flash
The first clue was timing. Google had already released a model under the Gemini 3.8 Flash name, so seeing another Arena entry with exactly the same label was unusual.
The second clue was capability. Developers testing the new variant reported results that felt closer to a frontier Pro model than a fast, efficiency-focused model. Complex instructions were followed in one pass. Interactive visuals arrived with fewer structural mistakes. Three-dimensional scenes looked more coherent. The model also appeared unusually comfortable generating complete experiences rather than isolated snippets of code.

A new Google model variant appeared under the gemini-3.8-flash label. Source: Arena Tracker screenshot.
For a community that had been waiting months for Google’s next major Pro release, the conclusion was irresistible: Gemini 4 Pro might already be in public testing.
Why the Appearance Caused So Much Excitement
Google’s previous major Pro update, Gemini 3.1 Pro, had arrived roughly seven months earlier. At Google I/O, Sundar Pichai previewed Gemini 3.5 Pro for a June launch, but the release slipped repeatedly. Reports later suggested that the 3.5 Pro project had been canceled after failing to create enough distance from the rapidly improving Flash line.
Attention then shifted to Gemini 4. Its pre-training evaluations were reportedly strong, and the model was said to be progressing through post-training. Against that backdrop, an unexpectedly capable Google model appearing under an old Flash label looked less like a routine update and more like the first checkpoint of the next flagship.
A Leaked Benchmark Table Put Gemini 4 Pro on Top
Then came the chart.
A benchmark table circulating online placed the suspected Gemini 4 Pro ahead of GPT-6 Astra, Claude Fable 5.1 and several other leading models across coding, agentic work, document understanding and computer use.
The headline results included:
| Benchmark | Task | Reported Gemini 4 Pro result |
| DeepSWE v1.1 | Long-horizon software engineering | 88.7% |
| GDPval-AA v2 | Real-world knowledge work | 2,064 Elo |
| Terminal-bench 2.1 | Agentic terminal coding | 95.3% |
| Terminal-bench 4.0 | General agent capabilities | 69.7% |
| OSWorld-2.0 | Agentic computer use | 86.8% |
| CharXiv Reasoning | Reasoning over complex charts | 94.7% |
| LABBench2 | Real-world biology research | 93.8% |

The benchmark table circulating online.
The same chart listed a price of $2.25 per million input tokens and $11.25 per million output tokens. If those numbers carry into a public product, Gemini 4 Pro would not only be competing on capability; it would also be positioned aggressively on price.
Other testers claimed access to a backend configuration with a 10-million-token input limit, a 256,000-token output limit, persistent cross-session memory and direct web access without a separate API. Those claims added to the sense that Google may be preparing something much broader than a conventional chat model.
Even a Simple SVG Cat Became a Talking Point
One of the most shared comparisons was also one of the simplest: generate a cat as SVG.
The suspected Gemini 4 Pro produced a cleaner, more structurally consistent animal, while the Astra result looked heavier and less controlled. Nobody should choose a frontier model based on one cat, of course—but SVG is a surprisingly revealing test. The model must turn a visual concept into precise shapes, coordinates and layers while preserving proportion and style.

A community SVG comparison.
This was the first sign that Gemini 4 Pro might be especially strong at the intersection of code and visual design.
Web Design With Better Taste and Better Interaction
The larger demos made that impression much stronger.
Developer Bee reportedly asked the model to create a showcase website inspired by sketching, pencils and graphite. In roughly 14 minutes, it produced a page built around the idea of “scrolling as drawing”: lines became darker as the visitor moved down the page, turning a visual metaphor into an actual interaction.
Another test asked for a working site combining cyberpunk and retro-futurist design. The result included a 3D grid, a data dashboard and an interactive 3D object on the opening screen.
These examples mattered because the model was not simply choosing colors and placing cards. It was coordinating visual direction, layout, interaction and code in a single pass.
SVG and 3D Generation Became the Main Event
The most viral demo involved a pelican riding a bicycle.
The output included day and night modes, a working headlight, cadence controls, anatomical overlays and interactive elements. It looked less like a static generated picture and more like a miniature creative tool.


Gemini 4 Pro result showed a much more complete environment than Gemini 3.8 Flash.
Testers highlighted two improvements: the model was fast, and it could absorb a long list of requirements without dropping most of them. That combination is exactly what makes complex visual generation useful. A beautiful output is interesting; a beautiful output that already contains the requested controls and logic is a workflow.
Other examples pushed further into 3D:
- A pixel-art pagoda with an explorable sense of depth

- An Airbus H145 helicopter model produced in roughly ten minutes

- An interactive 3D flight simulation
- Small playable games with complete controls and visual logic

The difference in the flight-simulation test was especially striking. The public Gemini 3.8 Flash output looked rough and incomplete, while the Arena model produced a coherent mountain landscape, cockpit-style interface elements and a far more polished playable scene. That gap became one of the strongest arguments that the two identically named models were not actually the same system.
It Can Generate Playable Games in One Shot
Gemini 4 Pro also appeared comfortable building games rather than merely describing them.
Reported examples included a Minecraft-like world assembled from page structure through individual game modules, as well as a modern 3D kart racer generated in a short session. The interesting part was not just visual quality. The model had to create the environment, connect the controls, preserve game state and make the result playable.
That is a much harder target than a single image or webpage. It combines spatial reasoning, UI design, code generation, state management and performance—and it exposes mistakes immediately. If the early demos are representative, interactive generation may be one of Gemini 4 Pro’s defining strengths.
Is RSI the Reason Google Suddenly Looks Faster?
The other major theory surrounding Gemini 4 Pro is recursive self-improvement, or RSI.
The basic idea is that AI systems can help improve the process used to train, evaluate and refine the next generation of AI. Once that loop becomes productive, each model contributes to the development of its successor, potentially accelerating the entire research cycle.
Google DeepMind has already made RSI an explicit research topic. Its Dream-RSI work explores how agents can improve their own search strategies by learning from previous exploration. That does not mean a model is independently rewriting itself without human oversight. It does mean AI is becoming a more active participant in AI research.
The community theory is that Google has begun to close this loop internally—and that Gemini 4 Pro is the first visible result.
The Three-Way Frontier Race Is Back
The competitive landscape has changed dramatically. OpenAI has Astra. Anthropic has Fable 5.1. Both companies are pushing beyond chat into long-running tasks, coding, agents, reasoning and autonomous research.
For Google, simply releasing a model that is “better than Gemini 3.8” would not be enough. Gemini 4 Pro needs to demonstrate that the Gemini family still belongs at the frontier.
If the Arena model really is Gemini 4 Pro, the early signs are exactly what Google needed: standout visual coding, stronger 3D generation, fast execution and an ability to turn a dense prompt into a complete interactive result.
The model may be hiding behind a Flash label for now, but it has already succeeded at one thing: getting the entire AI community to pay attention again.