Featured Stories

How Trane gets building insights 60x faster with Amazon Bedrock AgentCore
The agent replaces a multi‑screen diagnostic workflow with a single natural‑language interaction, validated by internal ...
Read story →
How UK AISI and EvalEval Are Making Benchmark Results Reproducible
AISI released verified results for five benchmarks—HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench, and Termi...
Read story →

OpenAI launches GPT-6 Sol and Luna, boasting lower cost and fewer mistakes
The new Sol model targets complex tasks like coding, while Luna is aimed at high‑volume clerical work. API access costs ...
Read story →