Small Language Models And The New Economics Of Enterprise AI

Direct Source Verification: This story is aggregated from Forbes (forbes.com). Full reporting rights and copyright belong to the primary publisher.
Gary Kotovets, Chief Data, Analytics, and AI Officer at Dun & Bradstreet.

Gary Kotovets, Chief Data, Analytics, and AI Officer at Dun & Bradstreet.

getty​Enterprise AI is entering an accountability phase. Uber COO Andrew Macdonald ignited headlines when he said it was becoming harder to justify AI spending because higher token usage had not translated into a proportional increase in useful consumer features. His concern captures a question many companies now face as AI moves from pilots into production: How do we turn greater usage into sustainable enterprise value? ​

For many teams, the initial instinct was to send every problem to the largest model available. That approach helped companies experiment quickly, but production systems demand a different level of discipline. The question leaders should ask is: What is the smallest model that can do this job reliably?​

I think of small language models (SLMs) as smaller-scale operational units. Each one handles a discrete chunk of work, such as identifying a company website, classifying information, evaluating relationship signals or ranking relevant content. Because the assignment is narrow, the model can be trained and tuned to perform at scale and cost-effectively.​

Large language models (LLMs) still have an important role in complex reasoning and generation. SLMs can take on repetitive, high-volume work before an LLM is called, allowing the larger model to focus on the part of the process that requires its broader capabilities. The two primary benefits are straightforward: lower cost and stronger accuracy on a specific task.​​

Here’s what that looks like in practice: Consider a production workflow that creates structured company descriptions from content collected across approximately 500 million web pages. Sending all of that material directly to a general-purpose model would require it to process an enormous amount of irrelevant content along with the useful information.

A more efficient architecture is to use smaller models to chunk, filter and rank the source material before the LLM sees it. Only the passages most relevant to the company reach the larger model, while a validation layer checks the resulting description. In one such workflow at my company, this “filter before you generate” approach reduces the tokens sent to the LLM by approximately 95%.

Filtering can also help improve the input quality. More context does not automatically lead to a better answer, especially when much of that context is noise. A smaller, more relevant evidence set can improve the quality and consistency of the final output.​

At low volumes, token costs can appear manageable. But once a workflow must process tens or hundreds of millions of records, an architecture that worked during a pilot can become difficult to justify.​

Based on my own experience, I recommend evaluating model choice against the economics of the workflow before those costs become embedded in production. Larger models may be useful during experimentation, but if a task is repetitive, measurable and high-volume, it is worth testing whether a smaller model can handle it reliably. In many cases, the smaller model is not simply cheaper; it can also perform better because it has been optimized for a narrower job.​​

However, that focus comes with trade-offs. A task-specific model will not generalize across every problem, and a new capability may require fresh training data or a retrained model. Complex, open-ended and lower-volume work may still belong with an LLM. The goal is to match each part of the workflow to the model suited to it.​​

• Can it be divided into discrete tasks?

• Which parts are repetitive, measurable and high-volume?

• Could an SLM handle the filtering, classification or retrieval?

• Which decisions require broader reasoning, and when should the system escalate to a larger model or a human reviewer?​

Those questions should also be applied to systems already in place. Our team is beginning to reevaluate workflows built with LLMs to determine whether parts can be redesigned around smaller models. Importantly, being successful with this approach means allowing it to become part of how AI architecture is assessed, rather than viewing it as a one-time cost exercise.​

In the next phase of enterprise AI, I believe token consumption will be a weak proxy for progress. Useful, reliable outcomes are the measure that matters. Choosing the smallest model that can reliably perform each task is one practical way to accelerate value while keeping the economics sustainable.​​

Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

Original Source
https://www.forbes.com/councils/forbestechcouncil/2026/09/24/small-language-models-and-the-new-economics-of-enterprise-ai/
Visit Forbes ↗
SHARE STORY:
𝕏 f in

Related Coverage in Business