This is a deep dive into Artificial Intelligence (as it stands today), LLMs, Generative Search and the new autonomous ‘intelligent’ agents that will do our bidding. If you want to better understand the integration of technology in business and life in general this post is for you.
Without a doubt the news cycle in 2024 and every moment since November 30 2022 when ChatGPT was first released has breathlessly referenced some form of AI. From those who believe that we are about to be taken over by super-intelligent programs to those who think, a little more soberly, that the hype is over-selling AI much like we’ve been consistently oversold the possibility of the technological singularity: the moment machine intelligence surpasses that of humans.
For those of you who worry about this fearing that carbon-based lifeforms will lose their dominance of the planet and become obsolete, chill. The predictions on the singularity are based upon the iterative power of technology to scale its computing capacity and use data to better understand the world. While it’s true that Large Language Models (LLMs) trained on data exhibit remarkable reasoning skills they also have severe limitations imposed by the very nature of their architecture that constrain their ability to truly reason. They can no more ‘reason’ than a giraffe can fly, no matter how hard it may try.
Does that mean that AI is not going to happen? On the contrary.
Smart Automation is Baked-in Everywhere
We use the term “Artificial Intelligence” as a catch-all phrase to include every smart automation development. As data becomes more portable and devices become more powerful and easier to connect and talk to each other things appear to become more intelligent. My home patio lights come on automatically at sundown because a smart device connects to a smart hub that gets information about daylight hours from a weather service and knows exactly when the sun will rise and when it will set.
Security light settings kick-in at midnight and automated sensors track the movement of insects, pets and people separately and react to each one differently. None of that is inherently intelligent as we understand intelligence but all of it is smart and collectively it appears to be intelligent.
The ability of technology to create verticals that process large amounts of bounded data generates amazing value. In that regard technology works in exactly the same way technology has always worked: it augments human capability and shoulders the burden of repetitive or overly complex tasks so that humans can be more human and less machine-like. This is no different to the pothole machines replacing entire crews with picks, shovels and barrels except it now applies to brainwork.
To understand the impact this technology will have better consider that everything we do is a database activity. We, essentially, query a database to satisfy a particular need we have. When we ask our friends at a party which is the best bar to go to after midnight we are essentially querying a database that resides in the distributed knowledge found inside the heads of our friends and we act as the interface through which the query is shaped and the answer will be expressed.
We Query Databases To Get Answers
LLMs are exactly the same thing. They allow us to query a database and get an answer to satisfy whatever pressing need we have that is expressed through the query. Google search and the way it indexes an increasingly semantic web is no different.
At the party of our example the quality of the reply we are likely to get will depend upon the sum total of the knowledge and experience our friends possess. If one of them, for instance, is a barfly he’s likely to seriously raise the quality of the answer due to his specialized and, arguably, in-depth knowledge. If, however, all of our friends are teetotallers then the reply we’re likely to get back will be of much lower quality and possibly full of misinformation as they use what they have heard or can remember to fill-in the blacks and reply to our question, so they don’t disappoint us.
Unlike our friends LLMs are constrained by their training sets however. While generative search and generative AI appear to reason in essence they are simply copying the reasoning processes they have experienced in their training sets. In a bounded data set that is similar to what they have been trained with this is not an issue. But in an open data set the limitations of this approach become evident.
In a paper titled: GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models, Apple researchers cited specific examples and serious concerns regarding the ability of LLMs reason by explaining how they exhibited “Brittle Reasoning”. Even small, logical variations in a given context produce significant reasoning flaws because they break the reasoning patterns LLMs have been trained on.
This failing has been highlighted independently by neuroscience researchers intrigued by the way LLMs ‘learn’ and reason. LLMs do not learn general truths about the world, their research states, and the generative nature of their internal modelling is rife with improbabilities which helps explain the hallucinations they often experience. Indeed some researchers argue that hallucinations are a feature of an LLM, not a bug and can therefore not be ironed away by adding more parameters or training them on more data.
This doesn’t render LLMs useless but it seriously constraints their usefulness, particularly in cases where the logic of the query is not reflected in the logic of the LLMs training set. This is an important limitation to keep in mind, particularly as LLMs in the year ahead will have ever larger context windows that will enable them to appear more and more intelligent and therefore more useful.
A context window is the textual range around a token (which is part of a word) that an LLM can process when it generates information. Typically LLMs have faced significant constraints in the length of these tokens which is like a limitation on their memory and a constraint on their output. Google, amongst others, has found a way to overcome this difficulty through a summarising process that retains key aspects of the information being parsed so that LLMs can have infinite context windows.
Google’s former CEO waxes lyrical on the subject below (please add a pinch of salt in your own accord).
The advent of infinite context windows will make LLMs more responsive and make them sound more intelligent and possessive of better memory. These are important steps to more widespread use because they appear easier to use and more trustworthy than a human. We already know, from research, that humans trust robots more than they trust other humans. It’s not hard to imagine the same sense of trust is also extended to ‘intelligent’programs like LLMs who sound eminently knowledgeable and trustworthy.
That, however doesn’t make them useful. Not quite yet. Which is where generative search comes in.
Search Is Evolving
Search, of course, has always evolved. It’s evolved because human behavior (and the way we use tools like search) is in a constant state of flux. On the one hand, we want quick, accurate answers. On the other, content creators want to use every means available to game the system to their benefit. Both of these are human behaviors but their intent is incompatible.
Search is doing what it has always done: trying to reward the first subset with exactly what it wants and trying to stop the second subset from ever getting what it seeks. From attempts to turn Google into a veracity engine to constant algorithm updates designed to weed out undesirable or inaccurate content Google Search has always been less than perfect.
Google Search has always employed some form of AI in search but it is now going to have to do so in a way that allows it to also identify (and exclude?) AI-generated content that adds little or no value to a search query response.
IBM has a hybrid construct that uses a semantic search engine to feed an LLM in an attempt to weed out hallucinations and increase usefulness. In a hint of what’s to come a Data Science article shows how the implementation of an agent to a search engine and LLM can improve the performance of both (at least within the boundaries of the experiment).
Agents, of course, are not new. I first wrote about them back in 2014 and, since then, we have seen them fail with Amazon’s Alexa website first getting the chop in 2022 and then its voice assistant suffering some deprecation. Amazon wasn’t alone, Microsoft followed swiftly by ending support for Cortana, Apple deprecated Siri and Google did the same for Google Assistant.
Each of these companies had a different culture, outlook, history and intent when it came to voice search but they all, somehow, visualized the same thing: a seamless integration of a voice-powered digital assistant that would act like a virtual agent of sorts (albeit with severely limited autonomy) to carry out a host of human commands. The reality was somewhat different: most of us used voice assistants to turn the lights on and off, control ACs and ask about the weather.
Money Talks In Search
Voice search is expensive. It takes a lot to program a voice assistant, train it to recognize spoken words and then have it reply back in natural sounding language never mind the development, cost and maintenance of the ecosystem that grows around it. To give you an idea, Amazon reportedly invested $25 billion in Alexa, over 4 years and, over that time, never even came close to making a profit.
Not that LLMs are cheaper. The annual cost of running chatGPT is close to half a billion dollars. Google’s Gemini AI cost almost $200 million to develop and a rough calculation of its integration in search points to an annual cost running anywhere between $3 and $6 billion a year.
One of the reasons Google search became dominant was its cost-per-query ratio. By innovatively using shipping containers as their first data centers Google managed as early as 2003 to create a large watts-per-square-foot density that rendered it more competitive than any of its rivals at the time.
The current cost of LLMs is prohibitive for the long-term. The ease of use and widespread adoption because of that ease of use however, argue that cost-cutting innovations will be in the works sooner rather than later. This is where Agents come in.
The Future is Agentic
Imagine having a virtual butler that will actually perform tasks for you while you sleep: book concerts, arrange holidays, buy groceries and find out useful stuff you need. That, in essence, is an AI Agent. Computer scientists define an AI Agent as: “technological tools that can learn a lot about a given environment, and then – with a few simple prompts from a human – work to solve problems or perform specific tasks in that environment.”
Google has already announced Gemini 2.0, Microsoft is embedding Copilot everywhere, Apple has released Apple Intelligence across its entire Apple ecosystem and Amazon, still smarting from its Alexa foray, is sprinkling AI Agent fairy dust across its products.
While Large Language Models are created to understand and generate text and (increasingly images in multimodal setups), AI Agents handle tasks that require decision-making, real-world interactions, and autonomy. They also cost a lot less to develop (usually in the range of $50,000 to $150,000) and are a lot cheaper to run with their cost-per-query running to less than one tenth of that of chatGPT.
What makes them exciting however are two distinct capabilities: first the ability to query LLMs and take advantage of their text and image generative capacity and second their ability to learn from experience so that they become more personalized and productive in their environment.
Andrew Ng’s presentation on Agentic Reasoning is a little technical but if you bear with it, it is deeply revealing and totally exciting.
Whatever their simplicity or complexity AI Agents are made up of four distinct layers:
Input – A means through which we ask a question (if it’s a Knowledge Agent or give a command if it’s an Automation Agent). This is basically the command interface and it may range from voice to text.
Processing – This is the layer where all the important work is done as the Agent processes the data it needs and matches it to our command.
Action – This is the deliverable outcome.
Learning – The exciting thing about AI Agents is that they’re personalized. We all were so excited when Google Assistant learnt our accents and could understand anything from Queen’s English (a virtually obsolete form of spoken English) to a Glaswegian accent (a virtually incomprehensible accent to anyone who’s not from Glasgow). Imagine an Agent that learns what you like, how you say things or what you prefer by way of specific choices in a limited dataset of choices like winter holiday vacations, for instance.
An Agent that summarizes films, tells me which latest fantasy book I will find most exciting (and why) answers emails and goes looking for vegan recipes for me would free up a lot of time I spend performing all those tasks.
AI Agents For Business
AI Agents, of course, are intended for more than the frivolous pursuits I mentioned above. They can answer customer questions in conversational English and precise detail, they can optimize workflows by identifying and freeing up bottlenecks in information parsing or data processing.
Towards Data Science has a richly detailed piece on the future of Agents for business (and personal use) that includes the option for Agents to act autonomously and within the parameters that define their purpose, interact and collaborate with other autonomous Agents to bring about the desirable outcome.
Both Microsoft and Salesforce have invested heavily in this ecosystem so we are only seeing the beginning of Agent use but this much is clear: apps will slowly disappear much like the ten little blue links of search which we had to visit in turn trying to find the answer we sought and, failing that, we had to act as the refining factor and rephrase the search query. Agents will automate much of the dreary action of bringing up an app and drilling down to its function in order to perform something specific. We will, instead, ask our Agent (and individually as well as at business level we’re likely to have more than one) to perform specific tasks for us and achieve specific outcomes.
This means that SaaS will also go the way of the Dodo for much the same reason (though we may be able to rent instead of buy some Agents).
With so much automation happening around us two things will become evident: first, that despite our advanced technology we still used humans as the cheapest and easiest option to perform a machine’s work and second, that even the most intelligent humans perform some work that requires relatively little intelligence and can be safely delegated to a machine.
In this fashion we arrive at the final, pertinent question: what role for humans in a world where machines do a lot of the work and interact with and talk to other machines?
What Role For Humans?
The popularity of Microsoft Office and Microsoft Word in particular sounded the death knell of typing pools that took up entire floors of space and gainfully employed hundreds of human beings that were tasked to listen to Dictaphones and type up letters to be sent to their bosses for final approval and signature. Inconceivable as it may seem now that we employed people just to type up letters consider how we still use people to pick strawberries, though that too thankfully is changing.
Combining machine vision and robotics robot pickers are able to now perform the back-breaking work once done by humans and do it better, faster over a 21-hour period and more efficiently.
This now brings us to the true value of humans which, inevitably turns into a man-against-machine comparison. It is true that we all continue to perform tasks, each day, that don’t require particular intelligence and demand skills that can be broken down, turned into mathematics and programmatically fed into a machine. Examples abound from high-quality chefs
To baristas:
At the same time humans possess specific skills and knowledge of the real world that machines simply cannot grasp. At least not yet, if ever.
Take the example of a sniper. A human brain that has a good knowledge of the real world is trained to perform a synthesis of data that’s unique in each situation. To do that it synthesizes everything it knows with everything it is being fed and reaches decisions that allow it to outperform the most powerful supercomputer. While machines can synthesize variables quickly they cannot account for variations that change the image they’re processing on the fly. For that they need to start again, and again, and again in an endless loop that presents an insoluble problem better known in computer complexity theory as non-deterministic polynomial time completeness problem (NPC for short).
The human brain may process data much slower than the slowest computer today but it can perform such feats of computational agility because of its ability to cherry pick (in a cognitive version of the triage) what information to use in order to achieve a specific outcome.
If this is hard to comprehend consider that: “Caltech researchers have quantified the speed of human thought: a rate of 10 bits per second. However, our bodies’ sensory systems gather data about our environments at a rate of a billion bits per second, which is 100 million times faster than our thought processes.”
Intelligence, a brain-wide phenomenon, lies in the ability of the brain to pick just which 10 bits per second to process in order to get the job done.
To do that it uses intuition. Intuition is itself a construct made out of that brain’s constant “comparing [of] patterns of current environmental cues to stored patterns from previous experiences. The pattern matches are what provides you with intuition – or as it is sometimes framed – knowing without knowing how you know.”
This “knowing without knowing how you know” is what allows the sniper of our example, linked to above, to outperform not just machines but also the existing limitations of his own weapon. But if sniping is a little hard to wrap one’s head around there is another example we can draw on that perfectly exemplifies the power of machines to mimic human behavior and their limitations when it comes to intuition.
Machines Cannot Think
Formula one racing is really hard. It’s hard on the cars and it’s hard on the drivers. The variables around each circuit keep changing requiring amazing feats of concentration and decision-making. Yet a racing circuit is a strictly bounded environment that, in theory at least, machines should be able to crack. An autonomous vehicle driven by a powerful enough computer should be able to outperform a human.
The A2RL (Abu Dhabi Autonomous Racing League) tested this notion by replacing the human driver with 95kg of highly specialized and expensive computer hardware. The goal was to pit machines against humans with the former constantly learning and this improving in the bounded environment presented by the race track.
Initially, the performance gap between machine and humans was between three and five minutes per lap with the autonomous vehicle lagging behind. Over time this has closed to just eight seconds, an impressive feat but also an eternity in the world of sports car racing.
But what drove the limitations home was the carefully planned demonstration that was designed to showcase the progress made by the autonomous car against the human driver. In this case the autonomous car had a super-fast start but crashed into a wall within seconds of starting the race and was totaled.
A post-mortem analysis highlighted the fact that Formula One racing, much like most human activities that take us to the very edge of what we can do, require as much real-world knowledge as experience, awareness and skill on the racing track. The day of the race was a little cooler than usual. The human driver took this into account and compensated by rocking his car back and forth so that the tires could warm up and grip the race track. The autonomous vehicle, despite its computer having access to the same tire-temperature sensors, did not.
This was reasoning that was beyond the capabilities of its sophisticated hardware and software.
So while AI Agents will take over many of our most mundane tasks and increasingly sophisticated hardware and software will be able to perform work that we are currently doing, the true value of who we are will reside where it has always resided: in our ability to bring together real-world knowledge to solve problems where leaps of intuition are required.
The introduction of the Spinning Jenny in the 18th century did the work, per machine, that eight hand spinners could have done and kick-started a revolution in processes, scale and efficiency that is still running its course today. It forced us to reconsider the value of a human being and learn to train our brain to make the most of our unusual ability to synthesize complexity and find solutions machines cannot.
It is no different today.