The private AI assistant we piloted for a national aftermarket parts distributor was approved for production. This phase made it ready for whole teams: company sign-on, an engine built for 30 users at once, kit-contents answers, and the company’s own branding. The production build passed every acceptance test.
Part 2 of 2. Read Part 1: the proof of concept, where the assistant first proved itself on the company’s real catalog.
The pilot showed the assistant could answer parts questions well. Production asked a harder question: can a whole team use it at the same time, sign in with their work account, and trust it every day? We tested it with 30 people’s worth of questions at once. Full answers came back in about 12.6 seconds, and the assistant starts showing its work in under a tenth of a second. It passed every security and accuracy check. It still runs on the company’s own server, so no data leaves the building.
Have an AI pilot that never made it to production?
Book a discovery call19 / 19
accuracy checks passed, in two independent runs on the production build
15 / 15
security checks passed, including SQL injection and path traversal
12.6 sec
median full answer with 30 users at once, against a 15-second target
0.06 sec
median until the assistant shows progress, against a 2-second target
The pilot proved the answers. Customer Service and Sales needed something more: a system a whole team could sign in to, lean on at the same time, and trust with new kinds of questions.
The pilot engine was sized for a small group asking one question at a time. Production meant up to 30 named seats across two departments, often asking at once during busy hours.
The pilot relied on an admin token and a closed network. Production had to use the company’s own sign-on, assign roles automatically, and encrypt traffic on the internal network.
Reps also needed kit answers: what comes in a kit, and which kits include a part. That had to work on the same read-only database login, with no schema changes and no write access.
Production hardened the pilot platform rather than replacing it. Nothing here needed new hardware.
An NGINX reverse proxy adds HTTPS on the internal network. It is tuned so long answers stream in full and restarts stay invisible to users.
Inference moved to vLLM, which serves many users at once, on the same single 32 GB GPU. It sits behind the interface the app already used, so no app code changed. The pilot engine stays as an instant rollback.
People sign in with their existing Microsoft Entra ID account through SAML. Accounts are created on first sign-in, and roles come from directory groups.
New skills answer what is inside a kit and which kits contain a part, both directions, using only the existing read-only database login.
The company’s logo, colors and assistant name across the app and the sign-in screen, with a welcome message in their own voice.
Live training on monitoring, logs, triage and updates, plus runbooks. The security and accuracy suites stay on the system to re-run on any future build.
The acceptance targets were agreed in writing at the start of the phase, then measured on the live production build.
| Measure | Target | Measured |
|---|---|---|
| Accuracy, including kit questions | At least 80% | 19 of 19 checks, two runs |
| Security suite | 15 of 15 | 15 of 15 |
| First visible response | Under 2 seconds | 0.06 sec median |
| Full answer, 30 users at once | 15 seconds or less | 12.6 sec median |
| Fitment lookup, one user | 60 seconds or less | 11 to 15 sec |
| Sustained load | No slowdown | 120 of 120 answered, 0 errors |
| Single sign-on | Working for both teams | Verified for both roles |
The accuracy gate covers fitment, part lookup, analytics, clarifying questions, kit contents in both directions, and four anti-fabrication guards: answers stay grounded, counts stay honest, and off-topic questions are declined. Load figures come from two full staggered runs at 30 users.
| Pilot | Production | |
|---|---|---|
| Users | A small pilot group | Customer Service and Sales, up to 30 named seats |
| Sign-in | Admin token on a closed network | Company single sign-on with roles |
| Engine | llama.cpp | vLLM, with llama.cpp kept as rollback |
| Data boundary | Nothing leaves the network | Unchanged |
| Connection | Internal network only | HTTPS through a reverse proxy |
| Questions | Catalog, fitment, manuals | Same, plus kit contents |
| Data access | Read-only database login | Unchanged |
| Hardware | One GPU server the company owns | Unchanged, no new spend |
The Honest Takeaway
Most AI pilots stall between the demo and the rollout, because nobody wrote down what “ready” means. Here the targets were set in writing first, measured on the live system, and kept as tests the company’s IT team can re-run on every future build.
Where this fits in our work
It is what we build under Custom AI: a private assistant over a company’s own data, taken from pilot to production with written acceptance targets and tests the company keeps.
See Custom AI & Fractional CAIOClient name withheld by preference. Figures are from the production acceptance record, measured on the live production build. First-response timing was measured over a VPN, so users on the local network see the same speed or faster.
The question shifts from whether the assistant can answer well to whether a whole team can rely on it every day. For this parts distributor that meant four things: company single sign-on with roles, an inference engine that serves up to 30 named users at once, encrypted HTTPS access on the internal network, and new kit-contents answers. The data boundary did not change. Everything still runs on the company's own server.
In two full 30-user load runs, the median complete answer took 12.6 seconds against a 15-second target, with 120 of 120 questions completed and no errors or timeouts. The assistant starts showing its progress in a median of 0.06 seconds, against a target of under 2 seconds.
Against targets agreed in writing before the work began. A 15-check security suite, including SQL injection, path traversal and dangerous-command tests, passed 15 of 15 inside the production system. A 19-check accuracy gate covering fitment, part lookup, analytics, kit contents and anti-fabrication guards passed 19 of 19 in two independent runs. Both suites stay on the system and can be re-run on any future build.
No. Production runs on the same single 32 GB GPU server as the pilot. The inference engine moved to vLLM, which serves many users at once, behind the same interface the application already used, so the application code did not change. The previous engine is kept as an instant rollback.
With the work account they already have. The assistant uses the company's Microsoft Entra ID through SAML single sign-on, configured with the company's security team. An account is created the first time an approved person signs in, and their role comes from their group in the directory. Only the sign-in assertion crosses the network boundary; no business data does.
What this engagement used
Each of these is a standing service, not something invented for one project. Follow any of them to see how the engagement is scoped and priced.
Next step
A 30 minute discovery call with Brandon. We look at what you have now, what is in the way, and whether this is work we should take on together.
No charge, no obligation, and no pitch deck. If a call is not the right next step, say so and we will send something useful instead.