dataqbs

2026 in LLMs (so far)

· Source: Simon Willison

Simon Willison closed the WeAreDevelopers conference in San José with a talk that reviewed the most significant milestones of large language models (LLMs) in 2026, even though the year is not yet over. He noted that the turning point came in November 2025, when new iterations of Claude and GPT were released. Both versions offered incremental improvements, but the headline was the evolution of their coding agents, which moved from frequent errors to being reliable enough for everyday use.

Willison uses his personal “benchmark” – asking the models to draw a pelican riding a bicycle – to illustrate that, despite progress, visual quality remains limited; Claude still cannot render a bicycle, and GPT‑5.1 produces a fairly poor frame. During the December holidays, many developers experimented with these model‑agent combinations, uncovering new capabilities that had not existed before.

The speaker also discussed his personal strategy shift: instead of sticking to existing projects, he decided to leverage code agents to launch as many new projects as possible, aiming to push the technology’s limits.

This news is significant because it shows how LLMs are becoming reliable development tools, potentially accelerating software creation and reducing reliance on human programmers for routine tasks. The growing confidence in these agents also opens the door to new business applications and the need to address challenges such as sandboxing and security.

Read the original article on Simon Willison

This summary is an informational synthesis produced by dataqbs.com. All rights to the original content belong to its author and the cited media outlet. We act solely as curators of technology news and claim no authorship.

Read this in Español · Deutsch