AI Broke the Data Stack. Now What? My 7 Predictions for 2026
Seven Shifts Redefining the Data Stack in an AI-First World in 2026
As we step into 2026, I took some time to reflect on the last year in data. Every year feels like it moves faster than the last, but 2025 was something else entirely.
AI and LLMs have been around for years, but early last year, reasoning models dropped and changed everything. For the first time, rather than predicting the next token via pattern recognition, AI systems could follow step-by-step logic to solve complex problems.
Almost overnight, AI went from something fun to experiment with on the weekend to something that looked capable of taking on real business problems. Funding surged, enterprise adoption spiked, and vendors rushed to reshape their roadmaps around this new promise.
Just as quickly, the hype was cut short when a viral MIT report found that 95% of AI pilots are failing in production. We saw the same thing in our survey of 550+ data leaders this year: 84% of organizations said that they’re betting big on AI, but only 17% of them have actually made it a success.
The models are getting better and better, but they’re still falling short. What does this mean for the data world? Is the modern data stack failing, like so many people have said in the past? Or is it about to undergo a major evolution?
Based on research and conversations with hundreds of data leaders and practitioners, here are my seven predictions for the modern AI-native data stack in 2026.
But before we jump into the predictions, here’s my quick take:
If the last decade was about building the modern data stack, the next one will be about rebuilding it for a world where AI is not just a feature on top of data, but a major consumer, interpreter, and activator of it.
1. Analytics will be fundamentally reimagined around unstructured data
For decades, structured data has been at the center of analytics, with analysts approximating the messy real world via neat tables and metrics. We gained valuable insights on what happened, but lost nuance on the why.
This is an inevitable tradeoff since humans can’t process the entirety of every interaction, but AI can. Working with massive unstructured data isn’t a trivial task, but its value can’t be ignored. Research shows that 45% of companies cite “unstructured, fragmented data” as a challenge, yet the volume of business data grows 63% monthly on average, with up to 90% being unstructured.
There are signs that this technology is quickly maturing. Snowflake’s Cortex AISQL can handle declarative querying over mixed structured and unstructured data at scale. Microsoft, Google Cloud, and Salesforce have released similar features. Companies are already using AI to extract insights from knowledge bases, product feedback, and call center transcripts in ways previously impossible.
My prediction: Though metrics will never go away, analytics will be reimagined from the ground up, with unstructured data as a primary signal. As McKinsey noted, every enterprise that wants to be data-driven will soon need to “query and understand relationships between unstructured and semistructured data easier and faster.” This will fuel a broader change in rebuilding data pipelines for unstructured data and LLMs.
2. Data “agents” will go mainstream, freeing up time for higher impact work
Most data teams today are still burdened by manual grunt work. This leaves little capacity to operationalize AI, which is partly why 36% of our survey respondents said that their AI works in pilots but fails to gain real traction beyond that.
Data professionals are already testing the water with generative coding tools like Cursor to accelerate discovery, debugging, analysis, and support, even without formal agentic tooling in place. We’ve also seen signs of more complex use cases in action. Databricks research shows that LLM-generated data contracts can reduce manual effort by over 70%. Gartner predicted that by 2028, at least 15% of day-to-day work decisions will be made autonomously through agentic AI.
My prediction: We’ll shift from today’s AI-assisted workflows to more autonomous AI-augmented workflows, where data agents take on meaningful analytical and engineering tasks like root cause analysis, data tickets, and table restructuring. Teams who continue with “business as usual” will soon be overshadowed by those who learn how to hand off their manual work to AI and focus on higher value work.
A key issue, though, is what context do you need to feed these AI agents? Just like a new data analyst, there’s only so much that an AI agent can do without proper onboarding and training.
3. The AI analyst will finally make the self-service dream a reality
For years, the data world has dreamed of self-service analytics, but after decades of work we’re still left with low adoption and dashboard fatigue. Conversational analytics tools and Slack agents are gaining traction, but most are still limited to quick experiments, ad-hoc queries, and small reports.
The core challenge has always been data literacy: business users understand the business but struggle with data tools, while analytics systems understand schemas and metrics but not business language. Mature AI has the potential to bridge this gap.
Early implementations show promise. Enterprise customers building iterative loops that leverage contextual metadata have seen AI analyst response accuracy improve by 5x.
My prediction: AI that can understand both business language and data semantics will finally unlock meaningful self-service analytics. Whether this replaces traditional BI entirely is up in the air, but I think we’ll see the growth of conversational BI interfaces connected to agentic protocols.
The question of who will own this, though, is still open. We’ve already seen data warehouses (Snowflake Intelligence, Databricks AI/BI Genie), BI tools (Hex, Looker, PowerBI), AI companies, and newer startups making a play for this space.
4. The context layer will become foundational to enterprise AI
AI doesn’t know what it doesn’t know. And without context, that blind spot becomes its biggest risk. In our survey, 49% of AI failures were attributed to lack of business context, 45% cited hallucinations, and 26% pointed to lack of trust in AI output.
When data lacks context, AI can’t understand the specifics of a company’s industry, jargon, structure, and edge cases. Think of the thousands of unwritten rules at your organization that people just have to learn: when to escalate, what that specific executive means by “the latest revenue metric,” or what dashboard to never use.
Industry leaders have framed this problem as “context engineering.” As Andrej Karpathy explained:
“When in every industrial-strength LLM app, context engineering is the delicate art and science of filling the context window with just the right information for the next step.”
The context balancing problem is exacerbated by context rot, where Chroma found that across 18 LLMs, “model performance grows increasingly unreliable as input length grows”.
Companies have already started focusing on the semantics question. Palantir’s emphasis on semantics with Ontology has been key to its massive recent commercial momentum. The Open Semantic Interchange initiative, launched by Snowflake, Salesforce, dbt Labs, and others, aims to accelerate AI and BI adoption by moving fragmented data definitions into a universal semantic data framework. Microsoft announced Fabric IQ, a unified intelligence platform powered by semantic understanding and agentic AI.
My prediction: We’ll see context and semantics treated as infrastructure, resulting in the emergence of a shared, federated foundational layer for enterprises. This “context layer” will encode how a company thinks, decides, and acts. It will turn scattered, static knowledge into a single, reliable frame of reference. This will capture, maintain, and deliver context across all AI use cases and workflows through four key capabilities: automated context extraction, curated context products, human-in-the-loop feedback loops, and a unified context store that AI agents, BI tools, and humans can all query.
5. The separate data & AI stacks will fuse together into a new AI-native data stack
The modern data stack (ingest > storage > transform > BI) and the “AI stack” (models > agents > applications) have evolved somewhat separately. This showed up clearly in our survey: 34% of data leaders say AI is siloed with limited access rather than integrated into their stack.
However, there are signs that this is about to change. The Fivetran and dbt Labs merger signals a unification between deep transformation and modeling. A few months before that, Databricks championed a vision of unified data, AI, and governance in a “seamless, open ecosystem” at their Data + AI Summit 2025. We’ve mostly stopped arguing about warehouse vs. lakehouse, and aligned on open storage and cataloging that can be used by AI and data infrastructure alike, which are becoming increasingly intertwined as AI permeates traditional data tasks and data infrastructure becomes essential to AI.
As Matt Turck wrote in his latest MAD (Machine Learning, AI & Data) Landscape, we’re at the “end of an era,” marking a switch from the modern data stack’s unbundling to greater consolidation. “In effect, data infrastructure and AI infrastructure are collapsing into one plane; the seams are where value leaks.”
My prediction: The separate data and AI infrastructure will collapse into a single “AI-native data stack” where data and AI work together. This unified stack will be owned and operated by combined Data & AI Platform teams. We’ll see the merging and evolution of a few current data categories: data lineage into AI lineage, data catalogs and access controls into an AI context layer, workflow orchestration with prompt orchestration, and ETL pipelines with AI model and pipeline orchestration. We’ll also see the creation of some new AI-centered categories: agent observability, unstructured data quality and tagging, and evals.
6. Data teams must evolve—they’re the best placed people to solve the AI context problem
The rise of AI isn’t just a technical shift. As infrastructure changes, so too must the people powering it. The data skills that dominated the last decade are becoming increasingly automated and early career data tasks are already being replaced.
ChatGPT Enterprise has seen an 8x increase in weekly messages, 19x increase in structured workflows, and 320x increase in average reasoning token consumption per organization, mostly used for writing, coding, customer support, and data analysis. Research from ADP shows this is already starting to replace human execution for early career tasks.
Meanwhile, enterprises are struggling with talent. In our survey, 31% of organizations cited lack of AI talent and 23% cited unclear ownership of AI as blockers to scaling.
But here’s the paradox: data people are actually best positioned to thrive in this new AI world. Unlike software engineers who work in deterministic systems where a function either passes or fails, data teams have spent their entire careers working in nondeterministic environments. Two analysts can explore the same dataset and reach different conclusions, or two data scientists can follow the same inputs and generate entirely different models. Iteration, ambiguity, and back-and-forth experimentation are the norm among data people.
LLMs behave much more like analysts than like software. They are inherently nondeterministic, and getting a model to reliably answer the same question the same way requires training, semantic guidance, careful prompting, iterative refinement, and lots of exploration. This is second nature for data people, but foreign to deterministic disciplines.
My prediction: As AI takes on more grunt work, old data roles will increasingly become obsolete, but data people will not. Their natural strength in navigating nondeterminism makes them the most valuable people today to take on the challenge ahead: setting AI up for success with semantic training, context engineering, and unified observability.
Data practitioners will evolve into roles focused on context engineering, semantic training, and AI observability. For example, Data Analysts will likely evolve into Analytics Context Engineers, and Data Engineers into Data & AI Engineers.
7. Apache Iceberg will go mainstream and drive the data stack to fundamental interoperability
For years, the data warehouse has been the center of gravity for analytics. Once data lands in the warehouse, everything else follows: compute happens inside the warehouse, and all tools (ETL, BI, governance, active metadata) connect to it. That creates data gravity and vendor lock-in.
Apache Iceberg fundamentally changes this. An open-source table format that sits directly on top of cloud storage, Iceberg turns stored datasets into high-performance tables with ACID transactions, schema evolution, time-travel queries, and more. This means that regardless of where data is stored, any compute engine can query it with Iceberg, which shifts gravity from compute to storage.
In the last couple of years, Iceberg has become the “industry standard.” AWS, Google, Snowflake, and Databricks announced native support for Iceberg, including governance through catalogs. Microsoft adopted Iceberg to share data across Snowflake and Fabric, and Salesforce built “zero copy support” for open data lakes with Iceberg. Meanwhile, dozens of established startups and open-source projects like Cloudera, Dremio, DuckDB, Firebolt, and Trino use Iceberg as their native table format. Apache even hosted the first ever Iceberg Summit in 2025, with widespread industry support.
My prediction: As Iceberg becomes the dominant table format, we’ll see a push for fundamental interoperability of data, leading to new types of compute engines and business models. Lock-in will fade and choice will explode as data teams can bring their own compute to any data storage system.
This will also make governance, lineage, and context more important than ever before. With multiple engines writing to Iceberg tables, we’ll need stronger governance and lineage systems. That’s why it’s no surprise that 47% of organizations reported in our survey that they’re now investing in data governance, access control, and metadata management, making it the top investment in our survey.
Looking ahead
The gap between AI’s promise and its delivered value remains wide, but the path forward is becoming clearer. The AI-native data stack will not emerge from incremental upgrades. It requires rethinking how we capture, process, understand, and activate data in a world where AI is a first-class consumer.
Organizations that treat unstructured data as foundational, empower data agents with proper context, invest in semantic infrastructure, and enable their data teams to lead will move from pilots to production. Those that do not will continue to struggle, regardless of how advanced their models become.
The future is not about choosing between data and AI. It is about building the unified infrastructure where both can thrive. The organizations learning from 2025’s failures are already moving. The question is whether we are building for this new reality, or hoping yesterday’s stack will somehow stretch to meet it.
The data community is split on many of these predictions. Some believe unstructured data will fundamentally reshape analytics; others think it’s overblown. Some see data teams as uniquely positioned to lead AI; others worry they’ll be automated away.
Rather than declaring winners, it’s worth hearing multiple perspectives. Events like the Great Data Debate is a space for practitioners, builders, and leaders from the data & AI space like Cindy Howson(ThoughtSpot), Barry McCardel(Hex), Tristan Handy(dbt Labs), Karthik Ravindran(Microsoft), Bob Muglia(Snowflake), and Jaya Gupta(Foundation Capital) others on the ground, to share what they’re seeing in production environments, beyond the hyped narratives.
If you’re wrestling with these questions in your own organization, the discussion might be worth following. See who’s debating what:
Join “The Great Data Debate”
The Insight Index: Your Weekly Data & AI Digest
Top resources and recommended reads, carefully curated for you.
Is Agentic Metadata the Next Infrastructure Layer? — Bill Doerrfeld
From Big Data to Big Meaning(Podcast) — Jessica Talisman & Stewart Alsop
Why Enterprises Still Struggle to Scale AI — Hasit Trivedi
Where AI is headed in 2026 — Ashu Garg
How Do You Build a Context Graph — Jaya Gupta
The State Of LLMs 2025: Progress, Problems, and Predictions — Sebastian Raschka, PhD
Data Engineering in 2026: What Changes? — Ben Lorica
Why Ontology, Context Graphs, and Decision Traces Are the New AI Substrate — Bijit Ghosh
That’s all for this edition. Stay curious, keep exploring, and see you all in the next one!
About Metadata Weekly
Metadata Weekly isn’t just a newsletter. It’s shared community space where practitioners, builders, and thinkers come together to share stories, lessons, and ideas about what truly matters in the world of data and AI: trust, governance, context, discovery, and the human side of doing meaningful work.
Our goal is simple, to create a space that cuts through the noise and celebrates the people behind the amazing things that are happening in the data & AI domain.
Whether you’re solving messy problems, experimenting with AI, or figuring out how to make data more human, Metadata Weekly is your place to learn, reflect, and connect.
Got something on your mind? We’d love to hear from you. Hit Reply!





One prediction I'd add: the cost reckoning arrives for most orgs by Q3 2026. Agentic workflows trigger 10–20 LLM calls per task versus one for a standard chat interaction so token consumption compounds fast when you deploy at production scale. Nvidia and Microsoft both confirmed this year that enterprise AI compute costs currently exceed equivalent human labor costs at scale. The organizations that come out ahead will be the ones that automate what delivers clear, measurable value per task rather than what's technically possible. Wrote about this alongside the productivity amplification side both sides of the equation matter: https://www.nilayshah.me/ai-agents-are-shrinking-middle-why-they-make-1-expert-more-powerful-than-a-full-team/