
By Stephan Rabanser, Sayash Kapoor, Rishi Bommasani, Andrew Schwartz, Arvind Narayanan
At Google’s developer convention earlier this week, the corporate launched its newest mannequin, Gemini 3.5 Flash, alongside a brand new agent app, Antigravity 2.0. To showcase what this new agent setup is able to, Google claimed {that a} workforce of brokers had constructed a complete working system. The hassle reportedly required solely a single immediate, price solely about $900 in API charges, and was carried out by a number of dozen subagents working collectively.
Does this imply that complicated items of software program can now be constructed cheaply by AI? Not so quick:
The “single immediate” declare is deceptive. The weblog put up says the working system was constructed from a single immediate. However midway by means of the put up, Google discloses that the immediate “ended up being many 1000’s of strains” lengthy. What number of makes an attempt did it take to generate the immediate? How particular have been the directions to the agent? With out these important particulars, it’s arduous to know if the key sauce is a greater mannequin or simply extra effort put into prompting the mannequin. Furthermore, the run was carried out on a scaffold with specialised roles, delegation to subagents, and an agent to detect and stop dishonest. Within the launch put up, Google views the scaffold as a product function. However we don’t know whether or not the scaffold was overfit to this activity of constructing an working system from scratch, or whether or not it might carry out as effectively on different complicated software program engineering duties.
Google’s writeup shouldn’t be specific about what counted as human intervention. The put up mentions that the ultimate run to develop the working system required “no extra steerage or corrections from a human.” Nevertheless it doesn’t outline that normal. It describes infrastructure to kill and restart caught brokers. The put up mentions an earlier run through which the brokers appeared to cheat, after which the workforce added anti-cheating measures and re-ran the duty. Nevertheless it doesn’t report dry runs as a part of the methodology. Nor does it clearly say whether or not any brokers escalated to a human, whether or not the ultimate run required any guide restarts, approvals, or fixes, or what number of retries it took till the agent was profitable.
The writeup doesn’t report any try to research whether or not the brokers wrote the code from scratch or copied current code from the web. To Google’s credit score, the weblog put up notes that toy working methods are widespread undergraduate course tasks, and public implementations are straightforward to search out. The put up itself raises the priority that the agent may have regurgitated data relatively than constructing the working system from scratch. Nevertheless it didn’t deal with this concern—there was no similarity evaluation or log evaluation to examine if the agent copied current code. Even when there was no direct copying, writing an working system may be comparatively straightforward for brokers due to patterns memorized within the coaching information, so this doesn’t inform us a lot about brokers’ potential to create novel items of software program.
Google has not launched the prolonged immediate, the code the brokers wrote, or the logs from the run, which makes it unimaginable to independently consider the claims. Releasing the supply code or the agent logs may have allowed unbiased researchers to judge the standard of the artifacts and reply questions resembling whether or not the agent was copying current code. The weblog put up solely features a brief video documenting a snapshot of the event progress and the general narrative of the experiment.
However, the weblog put up does report the precise greenback quantity for constructing the working system ($916.92), alongside the entire token price range (a complete of two.6B tokens). These figures present helpful context, which we wish to credit score Google for. Lots of the evaluations we beforehand surveyed didn’t disclose price in any respect, which made their headline claims arduous to check with different evaluations.
Nonetheless, Google’s weblog put up is successfully a press launch. We acknowledge that it’s unrealistic to count on it to be scientifically rigorous. Evaluations like this one, that means a long-horizon real-world activity evaluated on a single run with the experimenter narrating what the agent did, have turn out to be widespread. Since a lot of them have been carried out by AI firms, it’s straightforward to dismiss your entire style as puffery.
However that will be a mistake. We consult with the rising paradigm as open-world evaluations, and we acknowledge this development in a current paper (and an accompanying weblog put up). Crucially, we argue that open-world evaluations require a brand new set of methodological norms. Executed proper, they will present a priceless perspective that benchmark-based analysis can’t.
Google’s experiment does add to the mounting proof that brokers or agent groups can autonomously or near-autonomously work on sure sorts of duties for very lengthy intervals of time, making progress with out getting caught or confused. As we argue in our paper, benchmark analysis is successfully unimaginable for this sort of activity for a lot of causes together with price. So it’s an thrilling time for unbiased evaluators from academia, nonprofits, and authorities to step in and supply the type of rigor and credibility to open-world evaluations which might be unlikely to be present in AI distributors’ personal claims.

