Tag: agentic AI

  • Agentic AI Is About to Leave the Screen

    Agentic AI Is About to Leave the Screen

    Why robotics may become the next operating layer for AI, and what changes when the physical world becomes programmable.

    Andreas's view

    My read: robotics is still being framed as hardware, when the more important shift is that AI is becoming an operating layer for physical work.

    The market tends to split into two shallow stories. One treats robots as factory equipment. The other treats humanoids as spectacle: impressive demos, big forecasts, uncertain timelines.

    I don't think either framing is enough. The more interesting story is that agentic AI gives robotics a new operating layer. Robots are not only getting better bodies. They are starting to get better ways to interpret context, plan actions and coordinate with digital systems.

    That changes the question. It is no longer only: what can AI answer? It becomes: what can AI do when it can perceive, decide and move?

    For leaders, the implication is practical: start mapping where physical work could become programmable. The strategic question is not "Should we buy robots?" It is where sensing, decision-making, workflow automation and safe machine execution could change cost, throughput, resilience or customer outcomes.

    From chatbot to operator

    The first wave of generative AI lived in a text box. It wrote, summarized, translated, coded and made knowledge work faster.

    The second wave is more ambitious. Agentic AI plans, checks, books, routes, escalates and triggers workflows. It turns AI from an interface into an operator.

    Robotics is where the operator model starts to touch the real world.

    If chatbots made AI visible, and agents make AI operational, robotics makes AI physical.

    This is a much bigger jump than the interface suggests. A chatbot operates in language. A software agent operates in digital systems. A robot operates in environments where physics, safety, maintenance, regulation and human trust all matter at the same time.

    That is why I would not start this discussion with humanoids. Humanoids are one form factor. The bigger story is physical AI: models, sensors, actuators, chips, batteries, simulation, edge computing, fleet software and enterprise workflows coming together.

    Robotics is not one market

    Robotics market stack showing industrial, service, medical, defense, consumer, and humanoid robot segments
    Robotics is not one market. It is a connected stack of industrial, service, medical, defense, consumer, and general-purpose systems.

    Robotics is already a real market, and it is much broader than the humanoid headlines.

    Industrial robots remain the established core: welding, assembly, painting, material handling, electronics, automotive and semiconductor manufacturing. Professional service robots cover logistics, warehouse automation, inspection, cleaning, hospitality, agriculture, construction and security. Medical and care robots include surgical systems, rehabilitation devices and hospital logistics.

    Defense and security robotics adds unmanned aerial, ground, surface and underwater systems, counter-drone capabilities, explosive ordnance disposal, reconnaissance, logistics and infrastructure protection. Consumer robots cover domestic devices such as vacuums and lawn robots. Humanoid and general-purpose robots sit on top of this stack as an early-stage category for environments built around human bodies.

    The data matters because it grounds the story. The International Federation of Robotics reported 542,000 industrial robot installations in 2024 and a global operational stock of 4.664 million units. Asia accounted for 74% of new deployments. China alone represented 54%.

    Service robotics is smaller and more fragmented, but it is moving. IFR's World Robotics 2025 service robot summary reported that worldwide sales of professional service robots grew 9% in 2024 to more than 199,000 units. Medical robots grew 91% to nearly 16,700 units.

    Market forecasts point in the same direction, even if the exact numbers should be treated carefully. Goldman Sachs sees the humanoid robot market reaching $38 billion by 2035. Morgan Stanley outlines a much larger long-term scenario: a potential $5 trillion humanoid market by 2050, including supply chains, repair, maintenance and support.

    Defense is one of the clearest signals that robotics is becoming a strategic technology segment, not only an automation category. Fortune Business Insights estimates the military robots market at $19.82 billion in 2025 and projects it to reach $42.90 billion by 2034. The exact number matters less than the direction: militaries are shifting from isolated unmanned platforms toward fleets, autonomy, sensing, secure communications and human-machine teaming.

    The point is not to believe every forecast. The point is that robotics is starting to look less like a hardware niche and more like a debate about who controls the operating layer of physical work.

    Why the cycle feels different now

    Robotics has had false dawns before. What makes this cycle worth watching is that several constraints are shifting at once.

    AI models are becoming more useful for perception, planning and adaptation. Google DeepMind describes Gemini Robotics as bringing AI agents into the physical world; Google's later Gemini Robotics-ER 1.6 work focuses on spatial logic, multi-view understanding, task planning and success detection.

    Simulation is improving too. Robots need data, but the physical world is expensive and slow. Synthetic environments, world models and simulation frameworks can compress training cycles. That is why NVIDIA's physical AI announcement matters: Jensen Huang called this a "ChatGPT moment for robotics" and framed physical AI as models that understand the real world, reason and plan actions.

    Enterprise demand is also clearer than before. Labor scarcity, warehouse complexity, aging populations, healthcare capacity, nearshoring and infrastructure build-out all create demand for automation that can work beyond perfectly structured factory cells.

    The market is not waiting for household humanoids. It is starting with work.

    It is also starting with security. The U.S. Department of Defense's Replicator initiative is built around all-domain attritable autonomous systems: lower-cost systems that can be fielded, updated and replaced faster than traditional platforms. NATO's DIANA Rapid Adoption Service recently awarded an R&D contract for undersea robotics and describes its role as helping Allies "move faster from identified capability need to real-world solutions." That is the defense version of the same physical AI thesis.

    Operator overseeing autonomous drone, ground, and undersea robotics systems for defense and security missions
    Defense robotics is shifting from isolated platforms toward autonomous fleets, sensing, secure communications, and human-machine teaming.

    Humanoids are the headline, not the whole story

    Humanoids matter because the world is built for people. Door handles, stairs, shelves, tools, kitchens, hospital rooms and factory aisles assume a human body.

    If robots can operate in those environments, the cost of automation changes. Companies may not need to redesign every workflow around a fixed machine. The machine could adapt to the workflow.

    That is the promise. It is also where the hype gets dangerous.

    Most useful robotics deployments will start where the economics are precise: structured tasks, high labor scarcity, safety risk, repetitive physical work, expensive downtime or environments where human work is hard to scale.

    Amazon is a useful case because it shows the less cinematic version of the future. The company says it has deployed its one millionth robot and introduced DeepFleet, a generative AI foundation model designed to coordinate robot movement across fulfillment centers. The stated goal is a 10% improvement in robot fleet travel efficiency.

    That is how physical AI will often arrive: not as a robot that looks like a person, but as a system-level improvement in throughput, cost, safety or resilience.

    The recent signal: capital is moving toward physical AI

    The last few weeks made the theme harder to dismiss.

    Germany's NEURA Robotics announced a Series C round of up to $1.4 billion in June 2026, backed by investors including NVIDIA, Amazon, Qualcomm, Bosch, Schaeffler, the European Investment Bank and Tether. NEURA founder David Reger put the strategic point plainly: "The future of AI will not only live on screens."

    OpenAI is also leaning into the theme. Sam Altman's 2026 roadmap says 2027 may bring robots that can do tasks in the real world. Separate reporting on OpenAI Robotics hiring is best read as a secondary signal, not the core proof point.

    This is more than a robotics startup cycle. It is a convergence of AI labs, cloud-scale compute, semiconductor platforms, industrial companies and capital markets.

    That matters for Europe. If physical AI becomes an industrial operating layer, Europe is not limited to being a regulator of someone else's platform. Its manufacturing base, robotics suppliers, automotive sector, industrial software, safety know-how and Mittelstand process expertise could become part of the stack, provided capital, compute, talent and adoption speed match the ambition.

    The operating model question

    The real question is not whether to buy robots. That is too narrow.

    The better question is: which parts of the operating model become programmable when AI can act in both digital and physical environments?

    In logistics, software agents may forecast demand, rebalance inventory and dispatch autonomous mobile robots. In healthcare, AI may coordinate patient logistics while robots move supplies or support clinical workflows. In manufacturing, physical AI may help factories adapt faster to product variation, quality issues or labor constraints.

    In defense, the question is even sharper. Autonomous systems can extend sensing, logistics, surveillance, electronic warfare and force protection into environments where human presence is dangerous or too slow. This does not remove the need for human judgment. It raises the standard for command, control, accountability, cyber resilience and rules of engagement.

    The value is not the robot in isolation. The value is the loop: sense the environment, interpret the situation, decide what should happen next, act safely, learn from the outcome.

    That loop is what makes robotics strategically interesting.

    It also makes it risky.

    Governance moves into the physical world

    Executives reviewing governance controls for supervised robotics and physical AI in an automated operations environment
    Physical AI will require governance models that cover permissions, audit trails, human override, safety, and accountable operations.

    Enterprises are still learning how to govern text-generating AI. Physical AI raises the bar.

    A weak chatbot answer can mislead a user. A poorly governed software agent can execute the wrong digital workflow. A poorly governed robot can damage equipment, block a line or create a safety incident in a regulated environment.

    That means physical AI needs a governance model before it scales.

    Who owns the robot's actions? What permissions does it have? What tasks require human approval? How are decisions logged? How is an incident reconstructed? Who updates the model? Who certifies safety after the model changes?

    These are not IT questions only. They are operating model questions.

    My expectation is that the companies that do this well will not describe the work as a robot deployment. They will describe it as a redesign of work: human judgment where ambiguity is high, machine execution where repetition and safety allow, and clear escalation when the system reaches its boundary.

    What I'm watching

    Four things will tell me whether this thesis is right.

    First, whether robotics deployments move from isolated machines to fleet-level operating systems.

    Second, whether AI labs and industrial companies build repeatable safety and governance patterns, not only better demos.

    Third, whether customers buy measurable outcomes rather than robots: lower downtime, faster fulfillment, safer operations, more resilient logistics or higher asset utilization.

    Fourth, whether the market starts valuing robotics companies as platform ecosystems rather than hardware manufacturers.

    AI is moving from language to action. From action to coordination. From coordination to physical work.

    The first wave lived on screens. The next one will increasingly show up in the world those screens were designed to manage. Physical AI is not just a device transition. It is an operating model transition.

    Sources and further reading

  • MIT Called It a Disenchanted Intern. METR Says Check the Growth Rate.

    MIT Called It a Disenchanted Intern. METR Says Check the Growth Rate.

    Something happened this week that I keep turning over.

    MIT published findings this month showing that when 41 AI models were tested across more than 11,000 real workplace tasks, the result was, in their words, like a “disenchanted intern” — hitting minimum benchmarks about 65% of the time, but never exceeding 50% success on tasks requiring genuinely superior-quality output. If you work in software, marketing, legal services, or knowledge work of any kind, that’s the snapshot.

    METR — a nonprofit focused on measuring AI capabilities — published a different kind of snapshot. Their metric is the “time horizon”: the maximum length of autonomous task a frontier AI can reliably complete. In 2019, the best AI could handle roughly a two-minute task without human intervention. By the end of 2025, that had grown to roughly an hour. The doubling time across that whole period: around seven months.

    METR’s January 2026 update tightened that number further. Post-2023, the best estimate for the doubling period is now 130 days — closer to four months.

    My read on this:

    The MIT study and the METR data aren’t in conflict. They’re measuring different things at different timescales. MIT is taking a photograph. METR is measuring the shutter speed. And the shutter speed is getting faster.

    I don’t think the “disenchanted intern” framing is wrong — it describes today accurately. What I’m less sure about is the assumption, implicit in most of the coverage I’ve read this week, that “today” is a stable state. An intern who gets twice as capable every four months is not the same resource at the end of the year as they are today.

    What I keep returning to is the gap between the current snapshot and the trajectory — and the opportunity that opens up in that gap. The MIT data is a photograph of now. The METR data is the shutter speed. Anyone building workflows, designing teams, or structuring how they work around AI capability today is working from a reference point that will be measurably out of date within a single planning cycle. That’s an opportunity signal at a scale and pace most planning assumptions don’t account for.

    Three things I’m watching:

    1. Where the doubling curve hits friction. Every exponential eventually meets a wall — physical limits, data constraints, regulatory friction. METR’s time-horizon metric is useful precisely because it measures real-world task completion, not synthetic benchmark scores. When the doubling cadence breaks, that will be the signal that the curve has met something real. I expect that to happen. I just don’t know when.

    2. Whether “minimally sufficient” matters or not. MIT’s 65% minimally sufficient rate sounds modest. But most enterprise workflows run on people who are minimally sufficient most of the time. The threshold isn’t excellence — it’s “acceptable at scale, around the clock, at near-zero marginal cost.” That bar is lower than it sounds, and closer than the headline number implies.

    3. The infrastructure spend as an access unlock. Alphabet, Meta, Microsoft, and Amazon are projected to spend nearly $700 billion combined on AI infrastructure in 2026 — roughly double what they spent last year. That capital isn’t just building capacity for the current snapshot. It’s funding the cost compression that makes the next several capability doublings broadly accessible. When the infrastructure matures, the cost floor drops — and the surface area for building on top of it expands with it.

    The disenchanted intern framing is apt today. My expectation is that it’s a better description of 2025 than it is of 2027.

    References

  • The pilot-to-production gap is an execution problem, not a model problem

    The pilot-to-production gap is an execution problem, not a model problem

    What was announced

    Through the week of February 9–15, 2026, the enterprise AI deployment story sharpened around a paradox: 95% of generative AI pilots still fail to reach production, yet 42% of enterprises now run agentic AI in production and 72% have agentic systems live in production or pilot. Microsoft’s February enterprise update reframed Copilot from “assistant” to “governance-first agent” capable of completing entire workflows. Oracle introduced Fusion Agentic Applications for finance, supply chain, and HR. OutSystems research released the same week reported that 94% of enterprises adopting agentic AI now flag agent sprawl as a primary concern.

    What it means

    The two statistics are not in conflict. They describe two different populations of organizations. The 95%-pilot-failure number describes how the average enterprise treats generative AI: a proof-of-concept budget, a small team, and a handoff to operations that never happens. The 42%-in-production number describes a smaller cohort that has done the operational work — governance, identity, runtime monitoring, rollback procedures, and explicit ownership of the agent fleet. The gap between the two cohorts is not technical. It is procedural.

    Microsoft’s “governance-first agent” framing acknowledges this directly. The next phase of enterprise AI is not better models. It is the operating discipline around models — who deploys them, who owns them when they misbehave, who pays for the inference, and how the organization rolls back a bad agent without disrupting downstream work. That is a CIO problem, not a CTO problem.

    Andreas’s view

    My read on this: the production cohort is pulling away from the pilot cohort, and the gap is widening every quarter. The companies in production are accumulating an operational learning curve — what governance looks like, how to staff agent operations, how to track agent behavior in production, how to compose agents into workflows without losing accountability. The companies still iterating on pilots are accumulating learnings about prompts and demos. Those are different skill sets and they compound at different rates.

    I don’t think the next 12 months reward the companies that pick the best model. They reward the companies that figured out how to operate any reasonable model at production scale, with controls, with monitoring, and with an explicit chain of accountability when an agent does the wrong thing. Agent sprawl is the leading indicator that the operations layer is missing — when 94% of practitioners flag it as a top concern, the conversation has moved past whether agents work and onto whether they are manageable.

    The way I see it: the clearest signal a board can get on where an organization actually stands is whether the CIO can produce a production agent inventory — by name, by owner, by usage volume, by incident count. If the question produces a list, the organization is in the production cohort. If it produces “we are still piloting,” it is in the failure cohort, and the strategic gap to peers will be visible in operating costs by mid-2027.

    Three things I’m watching

    Three things I’m watching:

    1. I’ll be watching whether companies can produce a named, owned, monitored agent inventory with rollback procedures on demand — that capability is the clearest proxy I have for whether a real agent operating model exists or not.
    2. The organizations that interest me are the ones shifting pilot evaluation from “did the demo work” to “did the agent ship to production with controls in place” — and backing that shift by defunding pilots that stay in demo mode past a fixed time-box.
    3. The question I’d be asking myself is whether a dedicated agent-operations lead — with explicit authority over the production fleet and seniority equivalent to the head of enterprise systems — is in place. Without single ownership, sprawl is the default outcome, and I expect that to show up clearly in incident and cost data over the next several quarters.

    References and related signals

  • When 88% of organizations have adopted AI, adoption stops being the question

    When 88% of organizations have adopted AI, adoption stops being the question

    What was announced

    The Stanford HAI 2026 AI Index landed in mid-January with a set of numbers that close out a debate. Organizational AI adoption reached 88% globally. Global corporate AI investment more than doubled in 2025 to $581.7 billion. Generative AI hit 53% population adoption within three years — faster than the personal computer or the internet. Four out of five university students now use generative AI as part of their coursework.

    What it means

    When adoption crosses the 80% line, the question of “should we adopt” becomes structurally uninteresting. Every relevant comparison group has already answered it. What remains is differentiation — and differentiation in a world of universal access is harder, not easier, than in a world of selective access. The strategic margin moves from access to integration depth, from licenses to workflow penetration, and from procurement decisions to operating-model decisions.

    The investment number is the more telling signal. $581.7 billion of corporate AI investment in a single year is a capital allocation that prices in a specific belief: that AI capability will compound at a rate that makes today’s spending the cheap option in retrospect. That belief either turns out to be correct, in which case the laggards face a permanent gap, or it overshoots, in which case the survivors of the correction still own infrastructure and skills the laggards do not.

    Andreas’s view

    My read on this: the AI Index numbers are not a celebration of momentum, they are a notice of obsolescence. Adoption was the entry-level metric — the one that let companies say “we are doing AI” without committing to anything that mattered. With 88% adoption, that metric is exhausted. The companies that conflate “we have AI deployed” with “we have an AI strategy” will be the ones surprised in 18 months when peers with the same headline adoption rate are operating at a fundamentally different unit-economics base.

    I don’t think the next two years will be about adopting more. They will be about routing work differently — deciding which functions become AI-native, which roles get redesigned, which middle-management layers compress, and which workflows get rebuilt from the ground up rather than augmented. The companies treating this as a tooling question will keep the org chart they had in 2024 and bolt assistants onto it. The companies treating it as a structural question will redesign for AI-native operations and harvest a different cost base.

    My expectation is that boards still reporting on adoption rates are measuring the wrong thing entirely. The number that matters is the percentage of work routed through AI-native processes versus AI-augmented legacy processes. Those are two different cost structures and two different competitive positions. The first is a step change. The second is a feature.

    Three things I’m watching

    1. I’ll be watching whether companies move away from adoption KPIs toward integration-depth KPIs — specifically, the percentage of revenue-generating workflows that are AI-native, not just AI-touched.
    2. The companies that stand out to me will be the ones that build the comparison the AI Index doesn’t make for them: how their spend per FTE on AI infrastructure and tooling stacks up against the 90th-percentile peer in their sector. If that number isn’t visible to leadership, it isn’t informing strategy.
    3. I’ll be watching whether organizations use the next 12 months as a workflow-redesign window rather than a tooling-procurement window. The structural opportunity narrows the moment competitors finish their redesign.

    References and related signals

  • The agentic year begins underprepared

    The agentic year begins underprepared

    The year opens with a measurable gap. McKinsey’s 2026 trust maturity survey, fielded in December and January, puts twenty-three percent of organizations into the scaling phase for agentic systems and thirty-nine percent into experimentation. The remaining majority — nearly two thirds — has not yet begun scaling AI across the enterprise. The capability frontier moved twelve to eighteen months faster than the operating models around it. That gap is no longer an experimentation question. It is the year’s defining strategic risk.

    The boards that close this gap first will not be using better models than their competitors. They will be running organizations that can metabolize what the models already do. The constraint is no longer technology. It is adoption — and adoption is a leadership problem.

    The shift is structural, not cyclical

    Agentic systems are not a new feature inside a familiar product. They are a new class of worker. They take a goal, decompose it into steps, hold state across those steps, call other tools, recover from errors, and return a completed unit of work. That changes what a job is, not how a job is done.

    The 2025 narrative — copilots, productivity boosts, ten percent uplift — is over. The 2026 question is harder. What units of work no longer require a human originator? What units of work now require a human reviewer instead of a human executor? Which decisions can be delegated to a system that explains its reasoning? The companies asking these questions on a Monday morning are reorganizing. The companies still benchmarking model accuracy are stalling.

    The shift is one-way. No board will vote in 2027 to remove agentic systems from a workflow they reduced from forty hours to four. The architectural choices made this year will compound.

    Diagram of one human silhouette passing a goal to a central node that branches into multiple task arrows
    Goal in, decomposition out, no human in the loop between.

    The role change has already happened on the ground

    Inside organizations that have actually shipped agentic systems, the role redefinition is happening informally, by individual contributors, ahead of any HR process. A senior analyst who used to write three reports a week now reviews twelve agent-drafted reports a week and signs off on the analysis. A staff engineer who used to write three pull requests a day now reviews fifteen agent-generated pull requests a day. An account manager who used to draft proposals now edits proposals the agent has built from CRM context.

    The work that survives is judgment, taste, accountability, and relationship. The work that does not survive is execution under specification. Job titles still describe the second category. Job content has already shifted to the first.

    First-line managers feel this most acutely. They were trained to manage humans doing execution work. They are now managing humans doing review work, who in turn are managing systems doing execution work. That is a different management discipline — closer to portfolio management of automated processes than to people management of execution teams.

    A figure at a desk with twelve document icons floating above, marking one of them
    Three reports a week became twelve reviews a week.

    The organizational consequence is delayering

    Span of control widens when the work below each manager becomes more automated and more reviewable. McKinsey’s parallel work on the state of organizations points in the same direction: companies that scale agentic systems also flatten by removing one to two layers of middle management. The economic logic is direct. Middle layers existed to translate strategy into execution and to coordinate the humans doing that execution. When the execution is increasingly handled by systems and the translation is increasingly handled by models, the layer is doing less.

    This is not the 2024 layoff cycle that hit individual contributors. This is a 2026 reorganization that compresses the manager-of-managers layer. It is structurally different and politically harder. The people most threatened by it are the people running the budget meetings about it.

    Organizations that resist the delayering will have a temporary cost advantage and a permanent decision-velocity disadvantage. Decision cycles compress when fewer humans need to be in the loop. The competitor who removed two layers will commit to a market move three weeks faster. Over a year, that compounds into a different market position.

    Two org-chart pyramids side by side, the right one flatter, with an arrow indicating compression
    The middle layer compresses, span of control widens.

    So what boards should do this quarter

    Two actions belong on the Q1 agenda. First, demand a workforce plan that names the units of work moving from human execution to human review, with a twelve-month horizon. Vague AI strategies are no longer acceptable as deliverables; the question is which jobs, which tasks, which review cadences, which accountability lines.

    Second, name an executive owner for the operating-model redesign — not for AI strategy as a separate track, but for the way the company will be organized around the systems it has already deployed. The CHRO and the COO are the natural owners. The CTO is not. The technology decision is downstream of the operating-model decision, and treating it as upstream is how organizations end up with sophisticated tools and a 2023 org chart.

    The year that just started will be measured by the gap between capability and operating model. The companies that close it first set the pace for the rest of the decade. The risk is not moving too fast. The risk is moving too late. Execution speed will separate leaders from followers.