Where AI stands today
Reasoning: uneven
Strong on many well-defined problems, but novel reasoning and reliable verification remain inconsistent.
Planning: brittle
Models can decompose goals, but long-horizon plans still drift when tools, people, and changing conditions enter the loop.
Learning: limited
In-context adaptation is useful. Durable learning across sessions still depends on external memory, evaluation, and retraining systems.
Creativity: capable
Models generate useful combinations across media, while originality, provenance, and judgment still require human scrutiny.
Adaptability: variable
Transfer between domains is improving, but unexpected situations and recovery from errors remain difficult.
Autonomy: gated
Current agents need permissions, checkpoints, and human review for consequential or extended workflows.
How we got here
2017: The transformer architecture is published
The foundation modern language models are built on. Attention changed what a single model could learn.
2020: GPT-3 shows few-shot learning
The first time one model handled very different tasks from a prompt alone. A turning point.
2022: ChatGPT reaches the mainstream
Not a research breakthrough on its own, but the moment the general public started paying attention.
2023: GPT-4 and multimodal models
Vision, reasoning and tool use in a single model. The gap to human performance on many tests narrowed visibly.
2024: Agents and reasoning models
Models that plan, use tools and run multi-step tasks. Still early, and still dependent on review for anything consequential.
2025: Benchmarks keep climbing
Benchmark validity and real-world reliability remain contested, especially when tasks, tools, and operating conditions change.
What AGI means to us
We define AGI as an AI system that can learn, reason, and perform at or above human level across any intellectual task, without needing task-specific training. That bar is high. Current systems are impressive but narrow.
One possibility is that broader intelligence emerges through an accumulation of capabilities rather than one breakthrough. Whether those capabilities become AGI remains uncertain.
In the meantime, we focus on reviewed workflows that can be evaluated against a concrete operating need, with explicit limits and human approval where consequences matter.
