GPT-6 Astra, Looped Transformers, and Hidden Reasoning
OpenAI’s GPT‑6 Astra uses looped transformers to deepen reasoning while keeping parameters low, which can obscure its chain‑of‑thought traces but still offers powerful graphics, coding, and computer‑use capabilities.

GPT‑6 Astra, Looped Transformers, and Hidden Reasoning
Quick take‑away
OpenAI’s new GPT‑6 Astra is a leap forward in reasoning, coding, and especially graphics, thanks in part to a looped‑transformer architecture that re‑uses the same layers many times. The model still produces explicit chains of thought, but the looping can make those traces harder to extract.
1. GPT‑6 Astra in a nutshell
Astra is the most capable model I’ve tested to date. Its strengths are:
- Graphics & rendering – It outperforms GPT‑5.6 on 3D‑rendering benchmarks and can even generate realistic animations.
- Coding & math – On standard coding benchmarks it scores near‑perfect, and it achieves a 99.9 % score on the ARC‑AGI‑3 logic‑puzzle benchmark.
- Computer‑use – Astra can interact with macOS GUIs via a reinforcement‑learning harness that feeds screenshots back to the model. This allows it to perform tasks like drawing in MS Paint or manipulating spreadsheets.
Despite these advances, Astra remains a reasoning model. It is trained with reinforcement learning with verifiable rewards (RLVR) and still generates intermediate reasoning traces, though those traces may be less obvious when the model loops its internal state.
2. What are looped transformers?
A looped transformer re‑applies the same stack of transformer blocks multiple times to the intermediate representation. Unlike simply adding more layers, the weights are shared across passes. The idea dates back to the Universal Transformer (2018) and has recently been adopted by Astra.
2.1 How the loop works
- Input tokenization – The prompt is tokenized and embedded.
- First pass – The token embeddings run through the standard transformer stack.
- Re‑entry – The output of the stack is fed back into the same stack, repeating the process n times.
- Final output – After the last pass, the final hidden states are projected to logits.
Because the same weights are used repeatedly, the model can refine its internal representation without increasing the parameter count. This can lead to deeper reasoning while keeping inference efficient.
2.2 Why it matters for reasoning traces
When a model loops over its own hidden states, the intermediate reasoning that would normally be exposed as a chain of thought becomes buried inside the recurrent passes. Extracting a clean, step‑by‑step trace requires either a special hook or a separate decoding pass. Consequently, the public API may not surface the chain of thought, giving the impression that the model is “hiding” its reasoning.
3. Recent research insights
Several papers published in the last month shed light on the practical effects of looped transformers:
- Depth vs. width trade‑off – Experiments show that a 12‑layer looped transformer with 8 passes can match the performance of a 48‑layer feed‑forward transformer on language modeling benchmarks.
- Interpretability – Researchers found that the hidden states after each pass converge to a stable representation, suggesting that the model is performing iterative refinement rather than generating new reasoning steps.
- Training stability – Looping introduces a form of implicit regularization, reducing over‑fitting on small datasets and improving generalization to out‑of‑distribution prompts.
These findings support the view that looped transformers are a powerful architectural tool, but they also highlight the need for new debugging and interpretability methods.
4. Practical implications for developers
- Prompt design – When using Astra, keep prompts concise. The model’s looping mechanism can handle longer contexts, but overly verbose prompts may still degrade performance.
- Tool integration – Astra’s computer‑use harness can be extended to other GUIs beyond macOS. The same looped‑transformer logic applies, so you can build custom agents that interact with any desktop application.
- Chain‑of‑thought extraction – If you need a transparent reasoning trace, consider enabling a debug mode that exposes the hidden states after each loop pass. This can be useful for compliance or educational purposes.
5. Looking ahead
Looped transformers represent a promising direction for scaling reasoning without exploding model size. As more teams adopt this technique, we can expect:
- Better generalization – Iterative refinement helps the model handle novel tasks.
- New interpretability tools – Researchers will develop methods to visualize the evolution of hidden states across passes.
- Hybrid architectures – Combining looping with sparse attention or mixture‑of‑experts could yield even more efficient models.
In short, GPT‑6 Astra’s looped‑transformer design is a step toward more capable, versatile language models that still retain a form of explicit reasoning, albeit in a more subtle form.
TL;DR
OpenAI’s GPT‑6 Astra uses looped transformers to deepen reasoning while keeping parameters low, which can obscure its chain‑of‑thought traces but still offers powerful graphics, coding, and computer‑use capabilities.
Source
Read original sourceWhy we picked this
New transformer research and hidden reasoning techniques—core AI content.