This shared build is unavailable. It may have been moved or removed. Choose a build from the showcase.
THE SHOWCASE/ REAL AGENT OUTPUT
Show me what it built.
Same prompts. Different models. Put the results next to each other and look closer.
THE EXPERIMENT / 001
Give it tools. Let it cook. Keep the receipts.
24 prompts. One finished build per run. Made for curious people with eyes.
0models in the mix
0builds to inspect
0 / 24prompts attempted
0independently scored
The builds /0prompts
0 results, grouped by the prompt that started them.
//
The next experiment starts here.
Give an agent a prompt. Let it build and test. Bring back what it made.
results/model_name/run-001/index.html
A different medium. Same curiosity. Give the agent the whole production. Watch what comes back.
The challenge set 24
One complete prompt. Tools from the start. Your move.
Model scores
Independent evaluations of the finished builds.
Compare matching tasks, tools and budgets. Scores give each scored task equal weight. Different task coverage is not a controlled ranking.
Numerical benchmarks matter. Research and training need guardrails, and a score is useful evidence. I use them. But equal scores do not mean equal experiences with the models behind them.
Look at this Deep-SWE snapshot. GPT-6 Astra [xhigh] and Gemini 3.8 Flash [high] both display 74%, with different costs, token usage, and step counts. In my hands, these models don't even feel like they're playing the same game.
Deep-SWE snapshot. Different effort settings; captured values, not a live ranking.Full image ↗
PUT YOUR HANDS ON IT
Meet the model through its work.
What I keep missing in benchmark tables is the chance to try what the model built. Open the app. Push the controls. Follow a workflow until something breaks. You start to recognize a model's style, its strengths, and the things it consistently neglects.
That's what I wanted to make easier with Trial.
01 / MAKE THE MODEL WORK
Hard enough to show the cracks.
I chose tasks with enough complexity to make the agent work for it. A couple of model generations ago, I would have budgeted multiple hours for these jobs.
In my own Astra runs so far, each has delivered a build in under an hour. That's what I observed in those runs; correctness still has to earn its own evidence.
02 / LET PEOPLE JUDGE
You already know where to start.
Does the physics look plausible? Is the game fun? Does the interface hold together? You can form a useful first impression without being a domain expert.
Then use the formal checks to see whether that impression survives the requirements. Both kinds of judgment matter.
Beyond the fly-through.
I like a good voxel world. Those Twitter fly-throughs can show real visual flair. But for learning how a model builds, I want something I can push back on.
Disturb the physics. Make decisions in a playable game. Play an instrument. Push an interface until it breaks. That's why I prefer these demos: they expose cause and effect, strengths, and failure modes you can actually reproduce.
Voxel apps can do all of that too. A pretty camera orbit just hasn't demonstrated it yet.
The Video lab takes that curiosity into another medium. Research, script, motion, voice, sound, and the final edit: give the agent the whole production. You know when a film loses you, when a visual makes an idea click, and when the narration gets ahead of the evidence. Start there.
Keep the benchmarks. Get your hands on the work.
01 / THE INPUT
Choose a challenge.
Copy a complete prompt from the library. Keep the evaluator directory separate.
Let the agent plan, write, execute, debug and test. Browser validation belongs in the run. The capstone also asks it to research the web and choose its own idea beyond the other 23 challenges.
Every HTML challenge delivers one self-contained index.html. Keep the original artifact intact. Metadata and screenshots are optional. Video experiments have their own briefs in the Video lab.
03 / THE COMPARISON
Put the builds together.
Refresh the gallery. Each prompt gets its own row of model outputs. Open a screenshot to inspect the build, or read the original prompt alongside it.
Screenshots show the surface. Open the app to inspect its interactions. Record independent checks in report.json; keep the agent's own logs in evidence/.
Bind local reports to the artifact hash. The public showcase publishes HTML and thumbnails; evaluation files are excluded from the website export.
Python 3.11+ runs the local gallery without dependencies. The gallery does not install or launch source projects. See the repository README and EVALUATION.md for the full workflow.