Skip to content
Artwork for Intellectually Curious
Intellectually Curious · August 6 · 6 min

AREX The AI That Never Stops Improving

The Beijing Academy of Artificial Intelligence developed AREX, a family of recursively self-improving agents designed for complex, deep research tasks. These agents operate using a bi-level loop system: an inner research loop gathers evidence while an outer self-improvement loop audits the results against specific constraints to refine the final answer. To manage long-horizon tasks, AREX utilizes an autonomous context-update tool that condenses interaction history into a compact state without losing critical verified findings. The training process involves agentic mid-training and reinforcement learning, with a specific focus on "key steps" where decisive evidence is found or errors are corrected. Available in both a dense 4B model (Turbo) and a 122B Mixture-of-Experts model (Base), the agents consistently outperform larger baselines on various reasoning and tool-use benchmarks. These models demonstrate that recursive verification and targeted refinement significantly enhance the reliability of AI-driven research. Note: This podcast was AI-generated, and sometimes AI can make mistakes. Please double-check any critical information. Sponsored by Embersilk LLC

0:00-6:12

transcript

No transcript — this publisher did not publish one.

show notes

The Beijing Academy of Artificial Intelligence developed AREX, a family of recursively self-improving agents designed for complex, deep research tasks. These agents operate using a bi-level loop system: an inner research loop gathers evidence while an outer self-improvement loop audits the results against specific constraints to refine the final answer. To manage long-horizon tasks, AREX utilizes an autonomous context-update tool that condenses interaction history into a compact state without losing critical verified findings. The training process involves agentic mid-training and reinforcement learning, with a specific focus on "key steps" where decisive evidence is found or errors are corrected. Available in both a dense 4B model (Turbo) and a 122B Mixture-of-Experts model (Base), the agents consistently outperform larger baselines on various reasoning and tool-use benchmarks. These models demonstrate that recursive verification and targeted refinement significantly enhance the reliability of AI-driven research.


Note:  This podcast was AI-generated, and sometimes AI can make mistakes.  Please double-check any critical information.

Sponsored by Embersilk LLC

links1