The first generation of AI applications has largely been built by adding systems. Start with an application and a database, then add an LLM. Add a vector database for semantic retrieval, an embedding service to populate it, a search engine for exact retrieval, a cache to control latency and cost, and an agent framework to coordinate everything. Background jobs, queues, schedulers, and observability often follow.
Each of those technologies may be excellent at what it does. The architectural problem is what happens between them. Every boundary introduces another network call, another credential, another deployment, another place for state to become inconsistent, and another system a team has to understand and operate.
As AI moves from experimentation into production applications, that complexity matters more. Intelligent applications increasingly need data, search, business logic, models, permissions, and application state within the same interaction. When each capability lives in a different system, even a relatively simple request can become a distributed-systems problem.
There is another way to think about the AI stack: instead of adding more infrastructure around the application, move more of those capabilities into the application runtime itself.
A common assembled AI stack
.png)
The individual technologies may all be excellent. The architectural cost comes from everything between them: network calls, credentials, duplicated state, synchronization, deployments, and operational ownership.
AI applications are becoming a loop
Strip away the terminology and many AI applications perform some version of the same process: find the right information, understand what it means, decide what should happen, take an action, and learn from the result.
A support application might retrieve account history and relevant documentation, interpret what a customer is asking, classify the request, trigger the appropriate workflow, and record the outcome. A commerce application might retrieve products and customer context, understand shopper intent, decide which experience to present, respond, and use the result to improve what happens next.
Today, those steps are often distributed across different technologies. Full-text search finds exact words. Vector search retrieves information based on semantic similarity. A model interprets the request. An orchestration framework determines which tools to call. Application code applies the business rules, while databases and caches maintain the state required throughout the process.
The result works, but the application spends a surprising amount of its time and money (infrastructure) moving context between systems.
What if more of that happened inside the runtime?
Harper has always approached application architecture from the opposite direction. Instead of treating the database, application server, cache, messaging layer, and APIs as separate systems by default, Harper brings those capabilities into a single process on a multi-threaded unified runtime.
Over the last several releases, that same architectural philosophy has increasingly extended into AI. This is the more important story behind Harper 5.0 through 5.3. The individual features matter, but together they point toward a different model for building AI-enabled applications.
The goal is not to create another AI framework or to replace the models developers already want to use. It is to reduce the amount of infrastructure required around those models.
The AI application on Harper
.png)
The model can still live elsewhere. What changes is everything around it. Retrieval, state, application logic, permissions, decisions, and orchestration can operate within the same runtime.
That distinction is important. The argument is not that every model should somehow run inside Harper. The opportunity is to collapse much of the infrastructure surrounding the model, so the application can reach its data, retrieval capabilities, business logic, and AI services without traversing a long chain of systems.
The pieces have been coming together
This direction did not appear suddenly in Harper 5.3. The architecture has been developing over several releases.
Harper 5.0 established much of the foundation. Harper became fully open source at the core, expanded its storage capabilities, and continued the move toward a runtime where data and application logic can operate together. For AI applications, that smaller surface is significant because both developers and agents have fewer infrastructure boundaries to understand.
Harper 5.1 made AI a direct capability of the runtime. Applications gained a common interface for generating text and embeddings across model providers. Embeddings could be generated automatically as data was written. Harper added a built-in MCP server so AI clients could discover and invoke application operations directly, and model-driven applications could execute tool-calling loops without introducing another orchestration service solely to connect those pieces.
Harper 5.2 focused heavily on making these architectures practical at production scale. Vector retrieval became more efficient, application-level filtering could participate directly in retrieval, and capabilities such as scheduling, secrets management, caching, and routing continued moving closer to the application itself.
With 5.3, two additional pieces make the broader direction easier to see: native full-text search and model-assisted decisions.

The releases matter less as isolated feature lists than as a progression: more of the intelligent application can operate within the runtime.
Search is not one problem
Vector search has become closely associated with AI because it can find information based on meaning rather than exact wording. That makes it particularly useful for retrieval-augmented generation and other applications where a user's language may not exactly match the underlying content.
But semantic similarity is not the answer to every search problem. Sometimes the exact words matter. A query for a product number, policy name, error message, contractual term, or specific phrase may be better served by full-text search.
Modern AI applications often need both. A useful way to think about the distinction is that full-text search helps an application find the words, while vector search helps it find the meaning. Harper can now perform both alongside the data and application logic they serve.
Find the words. Find the meaning. Apply judgment.

From generating answers to making controlled decisions
Most conversations about AI applications begin with generation: send a model some context and ask it to produce text. But many production workflows do not actually need another paragraph of generated language. They need a judgment.
Should this support request go to billing, refunds, or engineering? Does this new product belong in one category or another? Is this alert a duplicate, something worth investigating, or noise? In each case, the possible actions are already known. The application needs the model to help choose among them.
This pattern has recently gained more attention through Jev and other decision-model approaches. Jev is a useful reference because it makes the distinction easy to see: instead of asking a model to freely generate an answer, give it a bounded set of possibilities and ask it to judge which one best fits the input.
The underlying pattern, however, is broader than Jev or any one specialized model. Small inference models have performed classification and ranking tasks for years, and modern language models can perform the same kind of bounded judgment alongside their generative capabilities. What changes is how easily that judgment can become a first-class part of application logic.
Harper 5.3 exposes that pattern directly through models.decide().
An application provides the context and defines the allowed answers. Harper asks the configured model to evaluate those choices and returns both the selected value and a score for each permitted alternative. For example, an incoming request might be evaluated against billing, refund, or bug.
The important part is not simply that the model chooses billing. The application also receives information about the alternatives, which gives it something concrete to act on. A strong result might continue through an automated workflow. An uncertain result might be sent for human review. A separate “none of these” score can indicate that the choices themselves may not adequately describe the request.
The model contributes judgment, but the application retains control over what the possible actions are and what happens after the judgment is made.
That makes decisioning a useful middle ground between traditional deterministic software and open-ended generation. Business rules can continue handling the situations where the answer is known precisely. Generation can handle tasks where new language or content is required. Decision models handle the increasingly important space between them: situations where software knows the available actions but needs contextual judgment to choose among them.
And because models.decide() is an interface rather than a Jev-specific feature, the underlying model can change. A highly optimized decision model such as Jev could sit behind it. A local model could serve the decision. A general-purpose LLM could do the same job. The application's decision contract does not have to change with the model.
Harper can also persist those decisions and associate them with later outcomes. That creates a path toward evaluating how different models behave on an organization's actual workloads and establishing automation thresholds based on observed results rather than assuming that a model score corresponds directly to real-world accuracy.
From model input to application decision

Jev helped bring attention to decision models, but the broader pattern is more important: applications can ask models for bounded judgments without handing control of the workflow to the model.
Search finds evidence. Decisions turn it into action.
This is where the retrieval and decision capabilities start to fit together. Full-text search can retrieve information when the exact words matter. Vector search can retrieve information based on meaning. The resulting context can then inform a bounded decision about what the application should do next.
The model does not have to invent the workflow. It does not need authority over the application. It helps evaluate the evidence within a set of choices the application already understands.
That is a subtle difference, but potentially an important one as AI moves deeper into production systems. Search helps the application find the evidence. Decision models help interpret it. Application logic remains responsible for what happens next.
Fewer systems changes more than the architecture diagram
Infrastructure consolidation is often described in operational terms: fewer services, fewer vendors, and fewer things to deploy. Those benefits are real, but intelligent applications make the implications larger.
AI interactions frequently require context from several sources before the application can do anything useful. When data, retrieval, application logic, permissions, model access, and workflow coordination all sit behind separate service boundaries, that context constantly moves across the network. Each transition introduces latency, consumes resources, expands the security surface, and creates another potential failure point.
Bringing more of those capabilities together changes the economics of the request itself. Less data has to cross service boundaries, fewer systems need to coordinate the same operation, and fewer infrastructure components need to scale independently.
The fastest network call is still the one an application never needs to make.
This is not about replacing the model
Models are evolving rapidly, and organizations should be able to choose the providers and models that fit their requirements. Harper's architecture is designed around that assumption.
The model is a configurable capability used by the application. The application owns its data, state, permissions, business logic, allowed actions, and operational behavior. Models contribute generation, interpretation, or judgment where those capabilities are useful.
That separation also makes the architecture less dependent on any particular model provider. As models improve or organizations change providers, the surrounding application does not necessarily need to be redesigned with them.
The AI stack may be starting to consolidate
New technology categories often expand before they consolidate. The first generation of a new architecture naturally produces specialized systems for every emerging problem because those problems need to be solved quickly.
AI followed that pattern. Teams assembled databases, vector stores, search engines, embedding pipelines, agent frameworks, model gateways, caches, queues, schedulers, and application servers into increasingly capable systems.
The next phase may look different. Search does not necessarily need to be separated from the data being searched. Embeddings do not necessarily require a standalone pipeline. Agent tools do not necessarily need a separate gateway. Model access does not necessarily belong in custom integration code across every application. And model-assisted decisions do not necessarily need to occur somewhere outside the application responsible for acting on them.
Harper 5.3 is another step in that direction.
As AI becomes part of ordinary application behavior, the infrastructure supporting it can become part of the application runtime too.
The future of the AI stack may not be a bigger stack at all. It may be a more unified runtime.






.webp)


