September 28, 2026
GPT-6 Astra Completes Portal in 24 Hours After 3,336 Tool Calls
AI News

GPT-6 Astra Completes Portal in 24 Hours After 3,336 Tool Calls

Sep 7, 2026

OpenAI’s GPT-6 Astra has autonomously played Valve’s puzzle game Portal in an experiment of computer use that included 3,336 tool calls and nearly 24 hours of runtime. The run’s reported API equivalent cost was $571.18.

The experiment provides a rare demonstration of the capability of GPT-6 Astra to maintain a long chain of visual interpretation, planning, and computer actions. However, it should not be taken as evidence that the model can independently play arbitrary games in real time.

GPT-6 Astra is described by the company as the most capable model for computer use, browsing, software engineering, cybersecurity, science and professional work. The company says the model can perform complex multi-step tasks and can adapt if the requirements change.

How GPT-6 Astra Completed Portal

The Portal experiment was performed with a setup that enabled GPT-6 Astra to communicate with the game through external interfaces.

Instead of a person controlling the character, the model analyzed screenshots, tracked the character’s location, and chose the next action. The game was eventually finished after 3,336 tool calls.

This sets the experiment apart from a regular gaming benchmark. The setup was not about measuring how fast a human-like player could finish the game or reaction time, but rather about testing if an AI agent could repeatedly observe its environment, reason on what to do next and act on that decision.

And the process went on for nearly a whole day.

The $571 Cost Needs Some Context

The cost of the Portal experiment as reported was $571.18 in API usage.

That number is the run-generated usage, expressed in API-equivalent terms. This is not to be construed as the automatic amount that the experimenter paid out of his/her own pocket.

The cost remains instructive, however, because it highlights one of the difficulties of ever more autonomous artificial intelligence systems. But the economics of calling a powerful model over and over can add up, even if an agent can do a long task.

In practical computing applications, speed and cost are nearly as important as ability.

Portal Was Not Running in Real Time

A key thing to note about the experiment is that Portal was frozen while GPT-6 Astra analyzed the data and decided what to do next.

This distinction matters.

Normally, a human player reacts continuously to a changing game environment. In this experiment, the model may spend time understanding the scene and figuring out what to do next.

So beating Portal doesn’t mean that GPT-6 Astra can sit down and play a game exactly like a human player right now.

The experiment shows instead a different capability: the maintenance of a computer-use workflow over a very large number of sequential decisions.

Why Portal Is an Interesting Test

Valve’s Portal is particularly useful for this type of experiment because its puzzles require players to understand spatial relationships.

Players must determine where portals should be placed, understand the movement of objects and use momentum, timing and the environment to progress through test chambers.

An AI agent therefore needs more than a single correct answer. It needs to interpret the current state and determine which action could lead toward the next objective.

GPT-6 Astra’s ability to repeat that process thousands of times is the more significant part of the demonstration.

GPT-6 Astra Is Designed for Longer Computer Tasks

The Portal experiment comes shortly after OpenAI introduced GPT-6 Astra as a model designed to handle complex computer-based work.

OpenAI says Astra can fill out online forms, update customer records, conduct online research, create websites, perform frontend quality checks and troubleshoot software. It can also work with documents, spreadsheets and presentations while following templates and adapting to changing instructions.

OpenAI reports that Astra achieves higher computer-use performance in its OSWorld 2.0 latency simulations while taking about 47% less time per task than GPT-5.6 Sol in the comparison presented by the company.

The Portal experiment provides a more unusual real-world example of the same general capability.

Instead of completing a business workflow, the agent was placed inside a game environment and asked to maintain progress across thousands of interactions.

The Experiment Has Important Limitations

The result should be viewed as a demonstration rather than a definitive benchmark for autonomous gaming.

First, Portal is a game released in 2007 and has been extensively documented through guides, walkthroughs and other material. That means the model’s prior knowledge of the game’s mechanics could potentially have contributed to the result.

Second, the experiment did not involve continuous real-time gameplay because the game was paused while the model processed information.

Third, the experiment measures one specific setup. It does not establish that GPT-6 Astra can independently complete unfamiliar games with no prior knowledge, different interfaces or substantially different control requirements.

Those limitations do not make the experiment uninteresting. They simply define what can reasonably be concluded from it.

What the Portal Experiment Actually Shows

The clearest takeaway is the ability to maintain a long-running observe → reason → act cycle.

Traditional chatbots can provide an answer and stop. A computer-use agent has to keep track of what has already happened, interpret new information and decide what should happen next.

That becomes increasingly difficult as the number of steps grows.

In the Portal experiment, GPT-6 Astra continued that process through thousands of tool calls before reaching the end of the game.

This is also consistent with OpenAI’s broader positioning of Astra as a model built for multistep work rather than isolated question answering. Its API documentation describes support for asynchronous tool calling and mid-turn steering, allowing applications to continue work while tools execute and incorporate new instructions during a task.

Why This Matters Beyond Gaming

The most interesting applications may have little to do with video games.

The same basic computer-use loop can potentially be applied to software testing, research, data analysis, office workflows and other tasks that require many connected actions.

OpenAI says GPT-6 Astra can already work directly with specialized software in areas including scientific research and engineering.

That makes long-horizon reliability increasingly important. An agent that performs one action correctly is useful. An agent that can perform thousands of related actions without losing the original objective could be considerably more valuable.

The Portal experiment is therefore better understood as a demonstration of persistent autonomous computer use than simply an AI beating a video game.


KEY TAKEAWAYS

  • GPT-6 Astra completed Valve’s Portal in an autonomous computer-use experiment.
  • The experiment involved 3,336 tool calls and nearly 24 hours of runtime.
  • The reported API-equivalent usage was $571.18.
  • The game was paused while Astra processed screenshots and selected its next actions.
  • The experiment demonstrates long-running computer interaction, but it is not a definitive benchmark for real-time autonomous gaming.
  • OpenAI positions GPT-6 Astra more broadly as a model for complex computer use, coding, research and professional workflows.

FAQ

Did GPT-6 Astra really complete Portal?

Yes. An independent experiment reported that GPT-6 Astra completed the original Portal using an external computer-use setup. The run involved 3,336 tool calls.

How long did GPT-6 Astra take to complete Portal?

The experiment took nearly 24 hours of wall-clock time. The game was paused while the model processed information and planned its next actions.

How much did the Portal experiment cost?

The reported API-equivalent usage was $571.18. This represents the reported model usage cost rather than necessarily the experimenter’s direct out-of-pocket payment.

Was GPT-6 Astra playing Portal in real time?

No. The game was paused while the model processed screenshots and determined its next action. Therefore, the experiment should not be compared directly with conventional real-time human gameplay.

What does the Portal experiment prove?

It demonstrates that GPT-6 Astra can sustain a long sequence of computer-use actions involving visual interpretation, planning and execution. It does not prove that the model can independently solve every unfamiliar game or perform real-time gameplay at human speed.


CONCLUSION

GPT-6 Astra completing Portal is notable less because an AI finished an old video game and more because it maintained a long sequence of computer interactions until the task was complete.

The 3,336 tool calls and nearly 24-hour runtime show both the potential and current limitations of autonomous agents. The experiment demonstrates impressive persistence, while the paused gameplay and substantial reported usage cost show why speed, efficiency and reliability remain important challenges.

As AI systems move from answering questions to operating software, experiments like this provide an early look at what long-running computer-use agents can actually do.


Leave a Reply

Your email address will not be published. Required fields are marked *