What Previous Tech Revolutions Tell Us About The AI Data Problem
Carl D’Halluin, Chief Technology Officer at Datadobi.
gettyLess than 12 months after ChatGPT was announced, Sundar Pichai, Google and Alphabet’s CEO, appeared on 60 Minutes and said this: “I’ve always thought of AI as the most profound technology humanity is working on. More profound than fire or electricity or anything that we’ve done in the past.”
Whether, three years on, we all agree with that (if we ever did) is an interesting question. Of the many extraordinary claims made about AI during the initial hype cycle, Pichai’s claim has to be right up there. But is it valid? The advantage we have today is experience and the ability to contextualize AI against other “transformational” tech trends.
Clearly, the hype is unrelenting, and businesses have been prepared to bet the farm that Pichai and many other industry luminaries were right, Google included. According to a report from the BBC, parent company Alphabet “expects to spend as much as $205bn this year, mainly on AI projects and infrastructure.” As we all know, that’s just the tip of an enormous investment iceberg.
AI may feel different from other tech trends, but the gap between technical potential and practical value is actually pretty familiar.
Major technology shifts have often created new operational problems alongside their potential benefits. The technology itself has rarely been enough to deliver transformation per se, and ultimately, organizations have typically needed better ways to understand and manage the infrastructure and data beneath it.
Take virtualization, for example, which became a mainstream technology in the early 2000s. At a fundamental level, it uses software to create virtual versions of computing resources, allowing multiple operating systems and workloads to run independently on a single physical server.
On the face of it, that doesn’t sound anywhere near as exciting as AI, but it gave organizations much greater flexibility in how they used their physical servers, and that was, by any definition, transformational. By 2035, the virtualization market is expected to be worth over $35 billion.
But looking back, that compelling proposition also created a new problem, in that IT teams could no longer rely on a straightforward view of which applications were running on which physical servers.
For example, virtual machines could be created for short-term needs and then left running long after their original purpose had passed. Organizations [SL1] needed management tools that could show what was running across the environment and how resources were being used. Arguably, virtualization only began to deliver its full value once businesses could manage the more complex environment it had created.
Then there’s cloud computing, which, before the arrival of GenAI, was the poster child for transformational tech. Over the last 20 years or so, cloud has fundamentally changed IT economics by allowing resources and capacity to be added when needed, rather than through large up-front infrastructure investments.
New applications and services could be introduced more quickly because organizations no longer had to wait for physical infrastructure to be installed. It also made it relatively straightforward to move data into new platforms, at least compared to legacy methods.
For many, migration became the most urgent tech priority, especially where it promised lower costs or greater flexibility. Some went all in and moved as much of their infrastructure, services and data to one or more outsourced platforms as possible. Today, the market has matured, and cloud strategy is driven more by pragmatism than dogma.
With the benefit of hindsight, moving data did not automatically create a well-managed data estate. Information could be spread between on-premises systems and different cloud environments, often with copies remaining in the original location.
Organizations still needed to understand what data they held, who was responsible for it, whether it remained necessary and a host of other considerations. Many have yet to strike the right balance, but that’s another story.
They also needed to know whether sensitive information had been moved into the right environment and whether access remained appropriate. Cloud providers could manage the underlying infrastructure, but they could not decide which data an organization should retain or how that data should be governed. Realizing the full benefits of cloud, particularly for those using it at scale, still depends on management and governance aligning with the speed of migration.
The list goes on, but whether we’re talking about the mobile tech revolution, the Internet of Things, big data, containerization, edge computing, APIs or a long list of other influential innovations, the technology created the opportunity first. In each case, organizations also had to work out how to manage the underlying data.
The same is true for AI, which depends on data an organization already holds, much of which has accumulated over many years across different systems and locations. A significant proportion of that information is likely to be unstructured, including documents, spreadsheets, presentations, images and video.
Organizations often lack a clear view of what this data contains, who owns it or whether it is still relevant. AI can make poorly understood data more valuable, but it can also make the consequences of poor data management more immediate.
The challenge is to understand the data estate well enough to decide what information AI should use, and do so well enough to help ensure the hype eventually matches reality.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

