Breakdowns of work on autonomous agents: planning, tool calls, multi-agent systems and why they fall apart on long tasks.
226 articlesTopics
The same posts, sorted by subject instead of date.
Models that write, read and fix code — from autocomplete to working through a task in a repository on their own.
39 articlesHow models arrive at an answer: chains of thought, planning, self-checking and what long deliberation costs.
29 articlesWhat a model remembers between requests and within one: long-term memory, long context, forgetting.
28 articlesModels that see: images, video, generation and visual understanding.
6 articlesHow model quality is measured and why benchmarks keep lying.
69 articlesHallucinations, alignment, robustness to attack, and what happens when models are trusted further than they should be.
9 articlesRLHF, GRPO, DPO and other ways to teach a model what the training data does not contain.
18 articlesModels driving physical devices: from manipulators to the gap between seeing and doing.
3 articlesText, sound, image and video in a single model.
6 articlesRetrieval-augmented generation and deep research: getting the right knowledge in front of the model at the right moment.
9 articlesDeployment, process change and economics: what happens to a company when AI stops being an experiment.
2 articlesData engineering in the model era: preparing data, analytics without SQL, and why text-to-SQL keeps breaking.
3 articlesConsciousness, AGI, trust, and what models do to the people who use them.
10 articlesInternal representations of reality: world models, simulation, and the move from prediction to understanding.
11 articlesModels beyond the screen: wearables, the internet of things and new ways to talk to a machine.
2 articlesGetting more out of a model: prompting, picking the right mode, and the mistakes people make using it.
29 articles