Skip to main content
Custom AI

The chatbot is just the tip: the nine layers under any AI a firm can trust

Brandon Aday, Founder of Aday Interactive, Inc.

By Brandon Aday

Founder, Aday Interactive, Inc. · Published September 24, 2026 · 9 min read

The short answer

A chatbot is the visible tenth of an AI system. Under it sit nine layers a firm never sees and cannot do without: identity and access, guardrails, agent orchestration, context and memory, retrieval, the model layer, tools and integration, the data foundation, and evaluation, observability and governance. A demo shows what AI can do; the layers make it do the right thing, repeatedly. The one number to watch is cost per outcome.

An iceberg: the chatbot above the waterline, nine layers of infrastructure below it

Every firm has now seen the demo. Someone types a question, the assistant answers in a warm paragraph, and the room decides this is the future. It is, and the demo is also the least important part of it. This article is built on Rathan Udayakumar's iceberg diagram, and it puts the point plainly. What the user sees is about a tenth of the system. The other nine tenths make an assistant safe to put in front of a client, a patient or a partner. They are where the real work lives, and the real money.

What does the user actually see?

A person, a chat window, an answer. Ask a question, summarize a document, draft a letter, pull an insight from a spreadsheet, kick off a task. It looks simple and feels magical, which is exactly the problem: the simplicity is an illusion produced by everything under the waterline.

The same interface can sit on top of a weekend prototype or on top of a system a regulated firm can defend in front of its licensing body. From the chair, they look identical. The difference is the nine layers below.

What are the nine layers under the chatbot?

The diagram names ten layers below the surface. This article folds observability into evaluation and governance, because for a firm of five to fifty people they are one conversation. Read them top to bottom, in the order a question travels.

  • Identity and access. Who is asking, what are they allowed to see, and is every request logged. Single sign-on, role-based access, audit trails. Without this, the assistant answers a paralegal with the partner's files.
  • Guardrails and security. The input is checked for prompt injection, the output for leaked data and off-limits content, the actions against policy. This is the layer that stops a client's chat from becoming a data breach.
  • Agent orchestration. Planning, reasoning, choosing a tool, taking an action, and asking a human before the action that matters. One assistant, or several that hand off to each other, with a person in the loop where the firm decides one belongs.
  • Context and memory. What the assistant remembers inside a conversation, across conversations, and what it is required to forget. Session management, relevance filtering, and retention rules that match the firm's own.
  • Retrieval and knowledge. The right document, from the sources this user is allowed to read, ranked well. Hybrid search (meaning plus keywords), reranking, permission-aware retrieval. The assistant is only as good as what it is allowed to find.
  • The model layer. The right model for the task: a large model for judgment, a small one for classification, a specialized one for the vertical. Routed by cost, speed, quality and compliance, with the option to run on your own cloud or on premises.
  • Tools and integration. The connectors to the systems where the work happens: the CRM, the practice management platform, the calendar, the billing system, the document store. Least-privilege access, so a tool can do its one job and nothing else.
  • The data foundation. Clean, current, classified, governed data. Ingestion, deduplication, metadata, chunking, embeddings, and lineage so every answer can be traced to its source.
  • Evaluation, observability and governance. Is the answer grounded in the source or invented. Did retrieval find the right thing. Did the tool do what it said. What did it cost. Can the firm show a regulator the policy, the audit trail and the explanation. This layer decides whether the system is trusted or merely used.

What is the difference between a prototype and an enterprise system?

The diagram's most useful panel is a two-column list, and it reads like a checklist of every AI pilot that stalled. A prototype has one model, a simple prompt, basic retrieval, open access to data, minimal security, no evaluation, manual deployment, accuracy measured on the demo, cost measured in tokens, and a chat window as the only interface.

An enterprise system has model routing and a prompt lifecycle (versioned, tested, retired). It has hybrid search with reranking, permission-aware data, guardrails and policy. It has continuous evaluation, an operations discipline for models and agents, and production reliability. It measures cost per outcome, and its agents take actions rather than only answering.

Notice that none of the right-hand column is about the model being smarter. Every item is about the surrounding system. That is the sentence in the diagram worth taping to the wall: the models are becoming one component, and the differentiation comes from everything around them.

Why does this matter for a firm of twelve people, not just an enterprise?

Because the risk does not scale down with headcount. A twelve-person law firm that lets an assistant read the document store has the same confidentiality obligation as a firm of twelve hundred. A concierge practice that lets an assistant see the schedule has the same patient-privacy obligation as a hospital. The layers are not enterprise vanity; they are the shape of the obligation.

What scales down is the implementation. A small firm does not need a platform team. It needs the same nine questions answered in proportion. Single sign-on it already pays for. One retrieval index over the documents it may expose. One model router. One evaluation set of fifty real questions with known-good answers. One policy document its people have actually read. Most of that is a week of careful work, not a year.

What does a real agent workflow look like?

The diagram's example is incident resolution, and the shape transfers to any firm's process: an alert comes in, the agent analyzes it, gathers context from the systems it is allowed to read, finds the root cause, proposes a fix, checks the fix against policy, gets a human's approval, executes, and validates the result before closing.

Swap the nouns and it is a new client inquiry. It arrives. The agent classifies it, pulls the intake history and drafts the response. It checks the draft against the firm's advertising and confidentiality rules, waits for the associate's approval, sends, and confirms the booking landed. The approval step is not a weakness in the automation. It is the point where the firm's judgment enters, and it is the step regulators will ask about first.

How do you know it is working?

Evaluation is the layer most pilots skip and the one that decides everything. The diagram lists eight measures, and a firm can run all of them on a spreadsheet before it runs them on a dashboard. Answer relevance. Groundedness, meaning the answer is supported by the source. Retrieval precision and recall. Hallucination rate. Tool-selection accuracy. Safety and policy compliance. Latency and cost. Human evaluation on a sample.

Observability is the same discipline for the running system: traces, logs, metrics, token usage, cost per task, agent steps, tool calls, and the business outcome each of them served. The last one is the only one the owner reads. If the system cannot say which outcome a cost served, it cannot be managed.

Where is your firm on the maturity journey?

The diagram draws six steps: chat (basic questions and answers), retrieval (the assistant reads the firm's own knowledge), copilot (it assists inside a workflow), agent (it performs a task), multi-agent (several coordinate across systems), and the agentic enterprise (continuous optimization). Most firms we meet are at step one and believe they are at step three, because the demo looked like a copilot.

The honest test is the layers. A chat with no retrieval is step one however it is dressed. A copilot with no evaluation set is a prototype however many people use it. Moving a step is not buying a better model; it is adding the next layer and proving it works.

  • Step one to two is a retrieval index over documents the firm is allowed to expose, with permission-aware access. This is where an AI Lab day starts: a receptionist and a qualification bot on the firm's own answers, in an account in the firm's name.
  • Step two to three is the copilot inside a real workflow, with an evaluation set and a human approval step. This is the Growth Systems Intensive and the 90-day Implementation.
  • Step three to four and beyond is an agent that takes actions across the firm's systems, with guardrails, tracing and cost per outcome. This is the Custom AI path: the Accelerator to prove one use case, Build Scoping to fix the scope, the build itself, and governance from the first commit.

The one number to watch: cost per outcome

Token counts are the prototype's number; they tell you what the model consumed. Cost per outcome is the enterprise number; it tells you what the firm paid for each thing it actually wanted: a qualified inquiry answered, a document summarized to a standard, a booking made, an incident closed. It folds in the model, the retrieval, the tool calls, the retries and the human's review time.

It is also the number that keeps the layers honest. A guardrail that blocks a bad answer lowers cost per outcome by removing the rework. An evaluation set lowers it by catching regressions before clients do. A router that sends simple questions to a small model lowers it directly. If a layer cannot be connected to that number, the firm is entitled to ask why it is there.

The takeaway, in five lines

The diagram closes with a sequence worth repeating. Model intelligence determines what the AI can do. The architecture determines what it should access. Policy determines what it is allowed to do. Evaluation determines whether it works. Governance determines whether the firm can trust it at scale.

A demo answers the first line. A firm that is going to put an assistant in front of clients needs the other four answered before it does. The ten questions below are how to ask.

FAQ

The ten questions to ask before choosing a model

Is our data ready, current and governed?

If the assistant will read the firm's documents, the documents have to be clean, deduplicated, classified by sensitivity, and current. An assistant over a messy document store gives confident answers from the wrong version. The data foundation is the first layer to build and the one most pilots skip.

Can retrieval respect permissions?

The assistant should only be able to find what the person asking is allowed to read. That means the retrieval index carries the same access rules as the document system, checked at query time, not a single index everyone can search.

Can every agent action be traced?

Every step an agent takes (the retrieval, the tool call, the draft, the send) should leave a trace that a human can read after the fact. Without it there is no audit trail, no way to debug a wrong action, and nothing to show a regulator.

Do we have guardrails for input, output and tools?

Three checks: the input for injection attempts and off-limits requests, the output for leaked data and prohibited content, and each tool for what it is allowed to do. A guardrail is a rule the firm wrote down, enforced by the system, not a hope that the model behaves.

Can we measure quality and detect hallucinations?

A set of real questions with known-good answers, run against the system on every change, with groundedness checked against the source. Fifty questions on a spreadsheet is enough to start; a system with no evaluation set is a prototype however many people use it.

Can we control cost per outcome?

The firm should be able to say what one completed outcome costs, including the model, the retrieval, the tool calls, retries and the human review, and see it move when a layer changes. Token counts alone cannot answer this.

Can humans intervene or stop the agent?

Yes, and the firm decides where. An approval step before any action that touches a client, a patient, money or a filing, and a stop switch that works. The human step is the point where the firm's judgment enters the process.

Can models be switched easily?

The model should be one component behind a router, so a better, cheaper or more compliant model can replace it without rebuilding the system. A system welded to one vendor's model is a prototype with a long contract.

Are we compliant with regulations?

The rules that already bind the firm (confidentiality, privacy, advertising, retention, the professional body's guidance) bind the assistant. Governance is the written policy, the audit trail, the explanation of how an answer was produced, and where the data lives. Nothing in this article is legal advice; the firm's own counsel decides what compliance requires.

Can we integrate with our core systems?

The value is in the connection to where the work happens: the CRM, the practice management platform, the calendar, billing, the document store. Each connection with the least privilege that does the job, and a way to turn it off.

Informational and educational purposes only

This article reflects Aday Interactive, Inc.'s views on marketing and technology architecture for professional-services firms as of the publication date. It is not a substitute for advice from a licensed professional in your jurisdiction and does not create any professional relationship between you and Aday Interactive, Inc. Rules, statutes, checklists, and AI-engine behavior referenced here can change; verify the current versions and consult qualified counsel before acting. Where the article discusses regulated professional practice, those references are for informational and educational purposes only and do not constitute legal, medical, tax, financial, or investment advice. Consult a licensed professional in your jurisdiction before acting on anything you read here.

Aday Interactive, Inc. provides custom web & SaaS development, AI search visibility (GEO/AEO/SEO), AI growth systems, and custom AI & fractional CAIO for established professional firms across the United States. Founder-led from Coral Gables, FL, with in-person engagements available throughout Miami-Dade County (Coral Gables, Brickell, Coconut Grove, South Miami) and remote delivery nationwide.

Free scorecard Request a consultation