Skip to main content
Custom AI · On-Premise · Production

From Pilot to Production, Still on Their Own Hardware

The private AI assistant we piloted for a national aftermarket parts distributor was approved for production. This phase made it ready for whole teams: company sign-on, an engine built for 30 users at once, kit-contents answers, and the company’s own branding. The production build passed every acceptance test.

Still Fully On-Premise Company Single Sign-On Load Tested at 30 Users
Explore Custom AI & CAIO

Part 2 of 2. Read Part 1: the proof of concept, where the assistant first proved itself on the company’s real catalog.

Marker sketch of a closed server cabinet, the production AI system running on the client's own hardware
In Plain Terms

The pilot showed the assistant could answer parts questions well. Production asked a harder question: can a whole team use it at the same time, sign in with their work account, and trust it every day? We tested it with 30 people’s worth of questions at once. Full answers came back in about 12.6 seconds, and the assistant starts showing its work in under a tenth of a second. It passed every security and accuracy check. It still runs on the company’s own server, so no data leaves the building.

Have an AI pilot that never made it to production?

Book a discovery call

19 / 19

accuracy checks passed, in two independent runs on the production build

15 / 15

security checks passed, including SQL injection and path traversal

12.6 sec

median full answer with 30 users at once, against a 15-second target

0.06 sec

median until the assistant shows progress, against a 2-second target

The Mandate

A Pilot Is Not a Production System

The pilot proved the answers. Customer Service and Sales needed something more: a system a whole team could sign in to, lean on at the same time, and trust with new kinds of questions.

Problem 01

Built for a Few Testers

The pilot engine was sized for a small group asking one question at a time. Production meant up to 30 named seats across two departments, often asking at once during busy hours.

Problem 02

Access Had to Match the Company

The pilot relied on an admin token and a closed network. Production had to use the company’s own sign-on, assign roles automatically, and encrypt traffic on the internal network.

Problem 03

New Questions, Same Rules

Reps also needed kit answers: what comes in a kit, and which kits include a part. That had to work on the same read-only database login, with no schema changes and no write access.

The Build

Six Workstreams, One Server

Production hardened the pilot platform rather than replacing it. Nothing here needed new hardware.

A Secure Front Door

An NGINX reverse proxy adds HTTPS on the internal network. It is tuned so long answers stream in full and restarts stay invisible to users.

An Engine Built for Teams

Inference moved to vLLM, which serves many users at once, on the same single 32 GB GPU. It sits behind the interface the app already used, so no app code changed. The pilot engine stays as an instant rollback.

Company Single Sign-On

People sign in with their existing Microsoft Entra ID account through SAML. Accounts are created on first sign-in, and roles come from directory groups.

Kit and Bill-of-Materials Answers

New skills answer what is inside a kit and which kits contain a part, both directions, using only the existing read-only database login.

Their Name on It

The company’s logo, colors and assistant name across the app and the sign-in screen, with a welcome message in their own voice.

Handoff to Their IT Team

Live training on monitoring, logs, triage and updates, plus runbooks. The security and accuracy suites stay on the system to re-run on any future build.

The Results

Every Target, Set Before We Started

The acceptance targets were agreed in writing at the start of the phase, then measured on the live production build.

MeasureTargetMeasured
Accuracy, including kit questions At least 80% 19 of 19 checks, two runs
Security suite 15 of 15 15 of 15
First visible response Under 2 seconds 0.06 sec median
Full answer, 30 users at once 15 seconds or less 12.6 sec median
Fitment lookup, one user 60 seconds or less 11 to 15 sec
Sustained load No slowdown 120 of 120 answered, 0 errors
Single sign-on Working for both teams Verified for both roles

The accuracy gate covers fitment, part lookup, analytics, clarifying questions, kit contents in both directions, and four anti-fabrication guards: answers stay grounded, counts stay honest, and off-topic questions are declined. Load figures come from two full staggered runs at 30 users.

Under the Hood

What Changed, and What Didn’t

PilotProduction
Users A small pilot group Customer Service and Sales, up to 30 named seats
Sign-in Admin token on a closed network Company single sign-on with roles
Engine llama.cpp vLLM, with llama.cpp kept as rollback
Data boundary Nothing leaves the network Unchanged
Connection Internal network only HTTPS through a reverse proxy
Questions Catalog, fitment, manuals Same, plus kit contents
Data access Read-only database login Unchanged
Hardware One GPU server the company owns Unchanged, no new spend

The Honest Takeaway

Most AI pilots stall between the demo and the rollout, because nobody wrote down what “ready” means. Here the targets were set in writing first, measured on the live system, and kept as tests the company’s IT team can re-run on every future build.

Where this fits in our work

This is a Knowledge Copilot, taken from pilot to production.

It is what we build under Custom AI: a private assistant over a company’s own data, taken from pilot to production with written acceptance targets and tests the company keeps.

See Custom AI & Fractional CAIO

Client name withheld by preference. Figures are from the production acceptance record, measured on the live production build. First-response timing was measured over a VPN, so users on the local network see the same speed or faster.

FAQ

FAQ: Taking On-Premise AI to Production

What changes when an on-premise AI pilot moves to production?

The question shifts from whether the assistant can answer well to whether a whole team can rely on it every day. For this parts distributor that meant four things: company single sign-on with roles, an inference engine that serves up to 30 named users at once, encrypted HTTPS access on the internal network, and new kit-contents answers. The data boundary did not change. Everything still runs on the company's own server.

How fast is the production assistant with 30 users at once?

In two full 30-user load runs, the median complete answer took 12.6 seconds against a 15-second target, with 120 of 120 questions completed and no errors or timeouts. The assistant starts showing its progress in a median of 0.06 seconds, against a target of under 2 seconds.

How was the production build tested before rollout?

Against targets agreed in writing before the work began. A 15-check security suite, including SQL injection, path traversal and dangerous-command tests, passed 15 of 15 inside the production system. A 19-check accuracy gate covering fitment, part lookup, analytics, kit contents and anti-fabrication guards passed 19 of 19 in two independent runs. Both suites stay on the system and can be re-run on any future build.

Did the move to production need new hardware?

No. Production runs on the same single 32 GB GPU server as the pilot. The inference engine moved to vLLM, which serves many users at once, behind the same interface the application already used, so the application code did not change. The previous engine is kept as an instant rollback.

How do employees sign in to an on-premise AI assistant?

With the work account they already have. The assistant uses the company's Microsoft Entra ID through SAML single sign-on, configured with the company's security team. An account is created the first time an approved person signs in, and their role comes from their group in the directory. Only the sign-in assertion crosses the network boundary; no business data does.

Next step

Let’s talk about your project

A 30 minute discovery call with Brandon. We look at what you have now, what is in the way, and whether this is work we should take on together.

  • A straight read on what is actually holding the work back
  • What an engagement would involve, in scope and in sequence
  • An honest answer if we are not the right fit

No charge, no obligation, and no pitch deck. If a call is not the right next step, say so and we will send something useful instead.

Part 1: Proof of Concept All Projects J. Randle Law
Free AI kit Request a consultation