A large on-premise custom AI deployment for a national aftermarket parts distributor. A five-figure catalog, vehicle fitment, and roughly 44,000 pages of service manuals answer plain-language questions in seconds, entirely inside the company’s own network.
91%
of benchmark questions answered correctly, against an 80% contractual bar
~11 sec
per fitment lookup, down from 3 to 8 minutes by hand
10k+ SKUs
plus ~44,000 service-manual passages, searchable from one assistant
100%
on-premise: zero catalog or customer data leaves their network
The expertise that made the business run was trapped in a catalog and thousands of manual pages, and the one tool that could unlock it, cloud AI, was off the table.
Product, fitment, and repair knowledge lived in a five-figure SKU catalog and thousands of manual pages. Answering one “what fits this vehicle, and how do I install it” question meant pulling a handful of veteran staff off their own work.
Off-the-shelf and cloud AI tools were a non-starter. The catalog, customer records, and proprietary manuals could not be sent to an outside service, retained by a vendor, or used to train someone else’s model. Whatever we built had to stay inside their walls.
A private assistant deployed as three services on a single GPU server they own: a local language model, a chat application carrying the agent logic and guardrails, and a vector search index over the manuals. It reads the live catalog through a read-only connection and cites the manuals for how-to answers. The source data never lands on the host.
Structural code guards close the failure modes that make AI risky in an operations setting. It declines out-of-scope questions cleanly, answers from the catalog and the manuals rather than from memory, and never invents a part number, a price, or a specification.
91% of benchmark questions answered correctly against an 80% contractual bar, scored on the client's own golden answer set.
A fitment lookup dropped from 3 to 8 minutes by hand to about 11 seconds end to end.
10,000-plus SKUs across 10 categories, structured fitment, and ~44,000 manual passages, searchable from a single assistant.
Model, index, and app all run on the client's hardware. No catalog, customer, or business data leaves the perimeter.
Fabrication, no-answer dodging, and internal-detail leaks closed with structural code guards, not just prompts.
A proof of concept on the original five-week scope, refined to nine weeks of testing and enhancements, then greenlit for production.
For an established company with a proprietary catalog and years of accumulated repair knowledge, the question is not only whether AI can answer well. It is whether the data can be trusted to an outside service at all. An on-premise deployment answers that plainly: every query, document, and embedding stays inside the company’s own perimeter. There is no vendor with a copy, no retention window to negotiate, and no chance the data trains a model someone else can use.
That single constraint, the data cannot leave, is what rules out most AI tools for firms like this one, and it is exactly the constraint a well-built on-premise system is designed around. The result is the capability of a modern assistant with the control of software you own outright.
The Honest Takeaway
On-premise AI is a real engineering build, not a subscription you switch on. What you get for that is worth it for the right company: the assistant runs on hardware you already own, the accuracy is measured against your own answer key, the guardrails refuse to guess, and not one byte of your business leaves the building. That combination is why this system was approved to move into production.
The whole system is version-locked into a signed container on the client’s own GPU server. It is auditable, restartable, and entirely theirs.
Local language model
A private LLM running on the client's GPU. It generates answers on the client's hardware, so prompts and responses never touch an outside API.
Chat application
The agent logic, the guardrails, and the read-only catalog connection. It routes each question to structured data or the manuals and enforces the rules that keep answers honest.
Vector search index
A local index over ~44,000 manual passages for how-to and specification questions, so the assistant cites the company's own documentation instead of guessing.
Where this fits in our work
It is the kind of engagement we build under Custom AI: a private assistant over a company’s own data, engineered for accuracy, guardrails, and full data sovereignty, and available with senior AI leadership through a Fractional Chief AI Officer when a firm wants the strategy alongside the build.
See Custom AI & Fractional CAIOClient name withheld by preference. Figures are from the pilot outcomes report on the deployed build.
An on-premise AI assistant runs entirely on hardware the company owns and controls, inside its own network. The language model, the search index, and the application all sit on the company's server, so questions are answered without sending any data to an outside cloud service. For this national parts distributor, a five-figure SKU catalog, vehicle fitment, and about 44,000 pages of service manuals became searchable in plain language, with nothing leaving the building.
Because the data cannot leave. When a catalog, customer records, or proprietary manuals are sensitive or competitively valuable, an on-premise deployment keeps every query, document, and embedding inside the company's own perimeter. It removes the vendor-access, data-retention, and train-on-your-data questions that make public cloud AI a non-starter for many established firms.
It answered 91 percent of the benchmark questions correctly against an 80 percent contractual bar. It is built to decline out-of-scope questions cleanly rather than guess, and it never invents a part number, a price, or a specification.
A vehicle fitment lookup that took a veteran staff member three to eight minutes by hand is answered in about eleven seconds. The product, fitment, and repair knowledge is no longer trapped with a few experienced people.
The working proof of concept landed on the original five-week scope. The engagement then ran to nine weeks as the client added testing rounds and enhancements, after which it was approved to move into production. A focused on-premise build is usually a matter of weeks, not months.
It runs as three services on a single GPU server the company owns: a local language model, a chat application with the agent logic and guardrails, and a vector search index over the manuals. It reads the live catalog through a read-only connection, and the source data never lands on the host. Access is least-privilege, named, and auditable.
Aday Interactive, Inc. provides custom web & SaaS development, AI search visibility (GEO/AEO/SEO), AI growth systems, and custom AI & fractional CAIO for established professional firms across the United States. Founder-led from Coral Gables, FL, with in-person engagements available throughout Miami-Dade County (Coral Gables, Brickell, Coconut Grove, South Miami) and remote delivery nationwide.