Building AI Systems That Reason Over Evidence, Not Assumptions
Why trustworthy enterprise AI should keep the LLM away from the database, calculator, integration layer and source of truth—and put evidence at the center of the architecture.
2026-10-04 · 9 min read · By Vikram C
The LLM should not be the system of truth
A language model is good at interpreting language, synthesizing context and explaining a conclusion. It is not a reliable database, calculator, integration layer or authorization engine.
For AEGIS, I am treating the model as a reasoning layer over evidence produced by tools and deterministic code. Source systems own the records. Code retrieves and validates them. Analytics code performs calculations. Authorization controls what can be accessed. The model explains what the evidence means.
Deterministic logic and probabilistic reasoning have different jobs
A completion rate should not change because a model interpreted a prompt differently. Neither should cycle time, review time or pipeline success rate. These are deterministic computations over defined inputs.
The model becomes useful after those facts exist. It can compare signals, summarize the context, identify a pattern worth investigating and explain the result in natural language. Keeping those responsibilities separate makes the system easier to evaluate and debug.
- Deterministic layer: retrieval, validation, calculations, permissions and state transitions.
- Probabilistic layer: interpretation, synthesis, summarization and natural-language explanation.
- Evidence layer: the records, metrics and provenance that connect the two.
A common data model is more important than a clever prompt
Enterprise APIs rarely expose information in the same shape. A project-management system might represent work items one way, while source control represents commits and merge requests differently and communication tools expose messages, threads and attachments.
A connector abstraction and common data model create a stable boundary. AEGIS can reason about concepts such as work items, changes, reviews, people, projects and evidence without coupling every downstream component to a vendor-specific payload.
Connector abstraction keeps integrations replaceable
Each connector should own the details of authentication, API calls, pagination, source-specific errors and translation into normalized entities. The agent should not need to know how every enterprise API works.
That separation also makes the model and client replaceable. A different model can consume the same evidence, and a different client can request the same underlying engineering context without rewriting the integration layer.
Evidence traceability changes the quality of an answer
An AI answer becomes more useful when the user can understand where it came from. Evidence can include source records, timestamps, metric definitions, connector results and the intermediate calculations used to reach a conclusion.
This does not mean exposing every internal implementation detail in every response. It means preserving enough provenance that important claims can be inspected rather than accepted because they sound plausible.
Security boundaries belong outside the model
A prompt saying “only use data this user can access” is not an authorization system. Identity, user scope and authorization need to be enforced before data reaches the reasoning layer.
In the current AEGIS prototype, Google Workspace authentication establishes the user context and the Gmail + ClickUp workflow is read-only. The architectural goal is to keep those permissions explicit and auditable as more connectors are added.
Cross-system reasoning is where the architecture becomes useful
Consider a question that needs both communication context and delivery context. A message in Gmail might explain why a task moved, while ClickUp contains the current task state and ownership. A useful answer needs both pieces of evidence.
AEGIS can retrieve the permitted records, normalize them, calculate any required metrics and then ask the model to explain the combined context. The model is not guessing what happened; it is reasoning over a deliberately assembled evidence set.
The architectural rule I keep coming back to
LLMs interpret and reason. Code retrieves, calculates, validates and enforces permissions. Enterprise systems remain the source of truth. Evidence connects the pieces.
That separation is the foundation I am using for AEGIS because it makes the system more explainable, more replaceable and easier to reason about as the number of connected enterprise systems grows.