AI progress is often reduced to a leaderboard: a model scores higher on reasoning, coding or multimodal benchmarks, and the industry moves on to the next release. But benchmarks are snapshots of what a model can demonstrate under controlled conditions. They do not always capture what happens when increasingly capable models are connected to an organization’s data, software, physical environments and decision-making processes.
The 2026 AI Index from Stanford HAI reports that frontier AI capabilities are advancing faster than many established benchmarks can measure, with some evaluations reaching saturation within months.
Members of the Senior Executive AI Think Tank see a similar pattern from the front lines of enterprise AI, machine learning, robotics, cybersecurity, cloud technology and digital product development. They point to a broader question for business leaders: What are today’s most capable AI models beginning to do that could matter far more than benchmark gains suggest? Here, they explore the emerging capabilities they believe deserve closer attention—and what those developments could mean for the way organizations build products, make decisions and operate in the years ahead.
“Everyone’s busy benchmarking models on text. The real shift is in how they read the physical world, and that’s the hard problem to solve.”
Bringing AI Into the Physical World
For Andrei Danescu, Co-Founder and CEO of Dexory, the next major frontier is not another text benchmark. It is perception.
“Everyone’s busy benchmarking models on text,” he says. “The real shift is in how they read the physical world, and that’s the hard problem to solve.”
Danescu cautions that robotics foundation models are still early, but he sees a significant transition underway.
“Robotics foundation models are claiming to transfer across robot bodies, generalize to objects they’ve never seen and reason before they act—but this is incredibly early days.”
The longer-term implication is bigger than a better robot interface.
“That won’t be a chatbot upgrade but rather a robot understanding its surroundings the way a model understands a sentence,” he says. “Atoms are next!”
Turning Enterprise History Into Institutional Intelligence
Aravind Nuthalapati, Cloud Technology Leader for Data and AI at Microsoft, focuses on a capability that is less visible but potentially just as consequential: AI learning from operational context.
“Frontier models can increasingly connect logs, conversations, documents, decisions, exceptions and outcomes across systems to reconstruct why something happened, not merely what happened,” he says.
That changes the role of enterprise AI. Instead of acting primarily as a knowledge assistant, a model could become an institutional intelligence layer that identifies hidden dependencies and patterns.
“The impact goes beyond automation,” Nuthalapati says. “Organizations could identify recurring failures, preserve expertise through workforce transitions and discover opportunities no one explicitly queried.”
This means leaders should treat operational history and decision context as strategic data assets, he says.
“The next AI advantage may come from organizational memory, not model size.”
From Answering Questions to Taking Action
David Obasiolu, AI Security, Governance and Systems Consultant at Vliso AI, sees the biggest change in what happens after a model generates an answer.
“The underrated shift is from answering to acting,” Obasiolu says. “Frontier models can now sustain multi-hour, multi-step work: navigating codebases, chaining tools, recovering from their own mistakes.”
That transition turns AI from a system people operate into a system that can operate software on their behalf.
Obasiolu sees two consequences. One is cybersecurity: “Models are finding and exploiting real vulnerabilities, compressing attacker timelines from weeks to hours.” The other is scientific acceleration, as AI begins helping researchers improve AI itself.
“The real story isn’t one dramatic capability,” he says. “It’s delegated autonomy spreading faster than our controls for it.”
Making AI a Scientific Collaborator
Matan Mishan, Senior Vice President, Agentic at Dot Compliance, brings an enterprise perspective to the question of where agentic systems could move beyond routine business workloads.
“Most people benchmark frontier models against today’s known workloads—tickets, code, summaries,” Mishan says. “That undersells what’s coming.”
He points instead to genome-driven drug discovery and longevity research, where AI-native companies are beginning to operate across entire scientific workflows: “The real unlock is models becoming genuine collaborators inside domains that used to demand years of specialized training.”
McKinsey’s analysis of life sciences workflows found that 75% to 85% of pharma workflows contain tasks that could be enhanced or automated by agents, illustrating why the gap between experimentation and operational adoption matters.
“Agentic AI is only now making its real debut in GxP-regulated R&D,” Mishan says. “That gap between ‘pilot’ and ‘new scientific method’ is where the real impact sits, and it’s opening this cycle, not in 10 years.”
“As capability becomes commoditized, the differentiator will be how intelligently—and responsibly—we integrate it into the way work actually gets done.”
Redesigning Work Around AI-Supported Execution
Aishwarya Shah, Independent Researcher, shifts the discussion away from individual model capabilities and toward the operating systems of companies themselves.
“The underestimated shift is not that AI can answer harder questions,” Shah says. “It is that models are becoming capable of sustaining context, reasoning across multiple forms of information and executing multi-step work.”
That combination can turn AI “from a tool we consult into a layer that can increasingly coordinate workflows.”
The consequence is organizational redesign. Processes built around handoffs, specialized knowledge and human coordination may increasingly be structured around AI-supported execution. But Shah argues that the advantage will not automatically belong to whoever has access to the newest model.
“The advantage will not simply go to organizations with the best models,” she says. “It will go to those that redesign workflows, governance and human accountability around them. As capability becomes commoditized, the differentiator will be how intelligently—and responsibly—we integrate it into the way work actually gets done.”
When AI Agents Become Negotiators
Bhubalan Mani, Leader in Supply Chain Technology and Analytics, focuses on an overlooked capability that is especially consequential for procurement: negotiation.
“Models can now bargain, trade off terms and close deals with other agents, not just draft the email for a human,” Mani says. Research into agent-to-agent commerce has found meaningful performance differences between models, including differences in the deals they secure and risks such as budget violations.
The implications extend to procurement. A buyer’s agent could negotiate simultaneously with thousands of suppliers over price, lead time and payment terms. “Contracts get renegotiated continuously, not annually,” Mani says, adding that model quality could quietly become margin.
“Few firms audit what their agent conceded or set walk-away limits it cannot cross,” he says. “In the agent economy, you won’t lose the negotiation. Your model will, and nobody will tell you.”
Using AI to See What Leaders Miss
Divya Parekh, Founder of executive coaching brand DivyaParekh.com, notes that her concern is not whether AI can generate more information but whether it can improve human judgment.
“What I think we are underestimating is AI’s ability to help us catch what we are not seeing,” Parekh says. Leaders already have more information than they can process, she adds, while important signals remain buried in noise.
The opportunity is for AI to challenge assumptions rather than simply confirm them. Frontier models are becoming better at connecting conversations, data, decisions and context, Parekh says.
“Imagine AI not just giving you an answer, but saying, ‘Here is what you may be missing before you make this decision,’” she says. “That starts to change the quality of judgment, not just productivity.”
The value moves from productivity toward decision quality—especially when AI is deliberately designed to surface counterarguments, dependencies and second-order effects.
“In a financial app, changing an automatic payment or finding a tax document can mean several menus. AI can understand the intent and take the user to the right place.”
Making Software Navigable Through Intent
Goran Paun, Principal and Creative Director at ArtVersion, believes one of the most consequential changes may look almost trivial: AI becoming a navigation layer inside everyday software.
“In a financial app, changing an automatic payment or finding a tax document can mean several menus,” he says. “AI can understand the intent and take the user to the right place.”
The same principle could apply to healthcare, HR and enterprise applications.
“The UI is not going away,” Paun says, “but AI may create a second path through software, built around intent rather than navigation.”
The significance is cumulative. Rather than replacing entire applications, AI could remove friction from thousands of interactions. Paun says these “small, indirect changes” can compound, saving time while reducing the resources required for routine tasks.
Finding the Questions Hidden in Enterprise Data
Maitrik Patel, Sr. Engineering Manager at Apple, argues that the strategic value of frontier models may lie less in answering known questions than in discovering unknown ones.
“The capability most underestimated is not what frontier models can do with a prompt,” Patel says. “It is what they can do with an organization’s own data that nobody thought to query.”
Years of support interactions, deployment logs, user behavior and internal decisions often remain disconnected because organizations first need to know which question to ask.
“Frontier models are beginning to close that loop,” he says. “They can surface patterns in data that predates the question, across systems that were never designed to talk to each other.”
The impact isn’t just fast answers, Patel adds, but better questions.
“The organizations that figure out how to point these models at their own institutional memory will have an advantage that no benchmark captures.”
Turning Dark Data Into Usable Intelligence
Vivek Kumkar, Sr. GenAI Leader at Amazon Web Services (AWS), highlights a category of enterprise information that has historically been difficult to search and analyze at scale.
“The capability most underestimated is that frontier models now read messy, real-world visual and multimodal content with close to the fluency they read text,” Kumkar says.
That includes video, scanned forms, screenshots, call recordings and images—information that companies possess but often cannot query efficiently.
“Most enterprise data was never text,” he says. Models are beginning to turn that material into “structured, searchable and actionable data in a single step.”
The implication is not simply better document processing. It could change what counts as an organization’s accessible knowledge base.
“The impact is not a better chatbot,” Kumkar says. “It is that the majority of a company’s information, the part that was effectively dark, becomes usable for the first time.”
What These AI Advances Mean for Business Leaders
- Look beyond text benchmarks and test AI in the physical world. Evaluate whether models can understand real environments, sensors, spatial relationships and changing conditions.
- Treat operational history as an AI asset. Connect decisions, exceptions, outcomes and workflows so models can identify patterns that traditional knowledge systems miss.
- Design for action, not just answers. As agents gain tool access and persistence, define permissions, monitoring, recovery mechanisms and human escalation before expanding autonomy.
- Put AI into domain-specific scientific workflows. Identify specialized processes where sustained context and multi-step reasoning could change how research itself is conducted.
- Redesign workflows around AI-supported execution. Do not simply insert a model into an existing process; reconsider handoffs, accountability and governance from the ground up.
- Govern agent-to-agent negotiations. Set explicit spending, pricing and concession limits, then audit what autonomous systems agree to on the organization’s behalf.
- Use AI to challenge decisions, not merely accelerate them. Configure systems to surface missing information, assumptions, counterarguments and second-order effects.
- Build intent-driven product experiences. Give customers a natural-language path through complex software while preserving traditional navigation for users who want it.
- Point models at institutional memory. Look for disconnected data sources where AI could uncover new questions, relationships or operational signals.
- Make multimodal information searchable. Treat video, audio, images, scans and screenshots as enterprise knowledge rather than passive archives.
The AI Opportunity Hiding in Plain Sight
The easiest way to measure AI progress has been to give a model a test and see what it gets right. But the business world is not a test. It is a warehouse, a clinical trial, a procurement negotiation, a messy database, a product interface, a meeting full of competing assumptions and millions of decisions that never make it into a benchmark. The capabilities that matter most may be the ones that allow AI to operate in that mess.
That creates a different challenge for executives. Instead of waiting for the next headline-making model, leaders can look at the parts of their organizations that remain difficult to see, connect, navigate or act on—and ask what happens when AI finally can. The biggest AI opportunity may not be something the industry announces but rather something sitting inside the business already, waiting for a model capable of finding it.
MOST POPULAR
9 Ways to Measure the Success of Your DEI Strategy
AI Is Commoditized—Here's What Sets Great Brands Apart
Inspiring Ideas. Actionable Insights.
Senior Executive's Email Newsletters Deliver Fresh Solutions to Today's Leadership Challenges.
Subscribe Free
Digital Assets: What TradFi Institutions Need to Know
Healthcare on Social Media: Building Trust in a Digital Age
From Clinical Complexity to Patient Action: A Better Content Model
