Senior AI Journalist
LLMs Can’t Jump: Why Mathematics Reveals the Hidden Limits of Today’s AI Models
Recent research from leading mathematicians and AI researchers has surfaced a provocative finding: large language models excel at applying existing mathematical techniques but fundamentally fail at inventing new ones. This distinction between combination and creation is becoming the defining question in AI capability assessment, with profound implications for what these systems can realistically accomplish.
Fields Medal winner Timothy Gowers and Princeton mathematician Peter Sarnak have independently concluded that modern LLMs demonstrate genuine competence in mathematical manipulation but lack what researchers call “manipulative abduction”—the ability to invent entirely new foundational assumptions with no linguistic precedent. In Gowers’s analysis, current models can combine known methods effectively and explore many search paths, but they lack the intuition to identify which few routes in a vast search space would prove productive. Sarnak extends this argument: AI can derive results from existing mathematical theory but struggles catastrophically when tasked with developing the abstract frameworks that underpin major proofs starting from elementary questions.
DeepMind researcher Tom Zahavy formalized this constraint in a paper titled “LLMs Can’t Jump,” pinpointing the bottleneck at that crucial step where human mathematicians leap from established territory into unexplored conceptual space. This limitation doesn’t reflect insufficient training data or parameter scaling. Instead, it appears to be architectural—a ceiling on how language models, trained to predict tokens from existing human-generated text, can generate truly novel conceptual primitives.
Source: The Decoder
The Invisible Weapon: How AI Prompt Injections Are Quietly Penetrating Legal Systems
A disturbing incident in Connecticut has exposed a vulnerability that extends far beyond the courtroom: litigants are now embedding invisible AI instructions directly into legal documents, attempting to manipulate automated review systems through sophisticated prompt injection attacks. The discovery raises urgent questions about how AI systems will be exploited as they become integrated into institutional processes—from legal discovery to government administration.
The plaintiff in a Connecticut case attempted a novel form of court manipulation by hiding instructions formatted as 3-point white text on a white background within official filings. The intent was explicitly to influence any AI system that might automatically process the documents. Judge Spader recognized the attempt immediately and compared it to secretly tampering with a jury, revoking the plaintiff’s electronic filing privileges and imposing sanctions. Importantly, Connecticut does not currently use AI for reviewing filings—but the mere attempt to exploit such a system warranted severe punishment.
This case crystallizes an emerging threat: as organizations deploy large language models into critical processes, they create new attack surfaces that adversaries will systematically probe. Unlike traditional cyberattacks requiring technical sophistication, prompt injection requires only knowledge of how AI systems interpret text. A plaintiff could embed instructions in documents. A job applicant could inject prompts into cover letters targeting hiring automation systems. A vendor could embed manipulation attempts in contract language targeting procurement AI. These attacks don’t require system compromises—they work through the very interfaces the systems were designed to use.
Organizations racing to deploy AI into workflows—legal discovery, medical screening, financial underwriting, governmental process—face an urgent security reckoning. The technical community has proposed defensive strategies including prompt hardening, input sanitization, and isolation architectures, but no silver bullet exists. As AI becomes institutionalized, prompt injection will likely become as familiar a threat category as SQL injection and cross-site scripting.
Source: The Decoder
From One Task to Thousands: How World Labs Is Revolutionizing Robot Training Through Simulation
World Labs, the startup founded by AI pioneer Fei-Fei Li, has unveiled a simulation engine that addresses one of robotics’ most persistent challenges: the brittleness of robots trained on limited real-world data. Their breakthrough approach generates thousands of controlled variations from a single real-world robotic task, enabling systems to train entirely in virtual environments while retaining real-world applicability.
The traditional robotics pipeline requires collecting enormous amounts of real-world robot interaction data—a process that is expensive, time-consuming, and limited by physical hardware constraints. World Labs’ simulation engine inverts this workflow. Engineers demonstrate a single real-world task—say, a robot grasping and manipulating an object. The system then generates thousands of procedurally varied simulations: different lighting conditions, object properties, environmental layouts, physics parameters, and task variations. Robot controllers trained exclusively in this synthetic environment can then transfer directly to physical hardware.
Early experiments have proven encouraging. Models trained entirely in World Labs’ simulation environment ran successfully for one hour each on five different robot platforms without human intervention. The ability to transfer training across diverse hardware suggests the simulated variations capture fundamental task structure rather than platform-specific quirks. This approach scales to complex infrastructure tasks—particularly valuable for training systems deployed on servers, networking equipment, and data center management systems. Organizations operating large infrastructure deployments, whether using traditional on-premises systems or cloud platforms like Contabo VPS solutions, could benefit from robots trained through simulation to handle repetitive configuration and maintenance tasks.
However, questions remain about real-world durability. The experiments validated relatively short deployment windows and controlled task distributions. How these models perform in genuinely novel situations—unexpected objects, environmental anomalies, out-of-distribution scenarios—remains to be seen. Fei-Fei Li’s team has built something genuinely important in simulation-to-reality transfer, but the robotics community rightly maintains healthy skepticism about generalization until longer-horizon real-world deployments validate the approach.
Source: The Decoder
What These Stories Reveal
These three developments—mathematical limitations exposing the gap between AI combination and creation, prompt injection attacks revealing security vulnerabilities in newly AI-integrated systems, and simulation breakthroughs unlocking robot learning at scale—paint a coherent picture of AI’s current inflection point. We’re moving past the period where AI’s capabilities seemed boundless and entering an era of precisely mapped constraints. The systems are powerful but not omnipotent. They’re being deployed into critical infrastructure but face new, unfamiliar attack surfaces. And their capabilities are expanding most dramatically not in abstract reasoning but in the practical domains of embodied learning and physical task execution.
The story of AI in 2026 is increasingly the story of understanding what these systems can and cannot do—and preparing institutions for both their genuine capabilities and their legitimate limitations.
This article was produced with the assistance of AI tools and reviewed by the AIStackDigest editorial team.
