Self-Hosted AI Layer for Customer Support

Location
USA, Czech Republic
Partnership period
2022 – Ongoing
Team size
5
Project information
Overview
A global infrastructure hosting provider operates data centers across North America, Europe, and Asia, serving customers from individual developers to enterprises running production workloads.
Its 24/7 support organization includes approximately 45 agents across three regions. The engagement began during a period of rapid growth, as ticket volume was increasing faster than headcount and the company was transforming its support operations across process, staffing, and tooling.

Challenge

Scaling support without compromising data control or response quality
As support operations expanded across regions, a significant part of the workload came before agents could start solving the actual customer issue. They had to reconstruct ticket history, search across fragmented internal knowledge, and coordinate work across shifts.
At the same time, tickets contained customer data and infrastructure details, which meant commercial LLM APIs did not meet the client’s requirements.
The key challenges were:
Complex cross-shift handoffs
Follow-the-sun support meant multiple agents could work on the same ticket, requiring each new agent to reconstruct its context.
Ticket volume outpacing headcount
Growing support demand increased the operational impact of time lost during repeated handoffs.
Fragmented internal knowledge
Answers were distributed across the knowledge base, billing platform, and internal notes, with limited search across them.
Inconsistent and error-prone responses
Response structure and completeness varied, while unsupported claims could lead to rework and escalation.
No safe path to public AI
Customer and infrastructure data could not be sent to commercial LLM APIs.
The client needed to reduce the time agents spent getting into a ticket, make internal knowledge answerable, and keep customer data entirely within the company’s own network.
Our Approach
Build AI around the support operation—not the other way around
Working closely with the client’s CTO, CodeIT designed a self-hosted AI layer around the existing support environment rather than introducing a standalone AI tool.
The architecture followed four principles from the start:
- Keep AI private. Generation and embeddings run inside the client’s VPN perimeter, with no customer data sent to third-party AI providers.
- Ground AI in company knowledge. Internal documentation and agent notes provide the evidence behind generated answers and fact-checking.
- Keep people in control. The assistant does not autonomously send customer responses or change, close, or assign tickets.
- Fit existing workflows. AI capabilities were embedded into the systems and team workflows agents already used rather than requiring them to move to a separate AI application.
This resulted in a five-component AI support layer running entirely within the client’s network perimeter.

The Solution
Ticket Analysis & Context Generation
CodeIT built a multi-stage enrichment pipeline that converts raw ticket text into structured support context.
The pipeline identifies signals such as urgency, politeness, and escalation pressure; extracts technical artifacts including IP addresses, ASNs, error codes, and server identifiers; and uses locally hosted models to generate structured summaries covering the customer’s request, work already performed, current blocker, and escalation risk.
Language resources were tuned specifically for hosting and infrastructure support rather than relying only on generic sentiment analysis.
Why it mattered: agents could work from structured ticket context instead of rebuilding the situation entirely from conversation history.

Knowledge Base Access Through the AI Assistant
The AI Assistant gives support agents conversational access to the client’s internal knowledge base, including years of runbooks, procedures, hardware guidance, and network documentation.
Agents can ask questions in natural language and receive answers based on relevant internal content, together with links to the supporting source articles. Behind the assistant, a RAG pipeline retrieves relevant information from the indexed knowledge base and grounds generated answers in company documentation.
Why it mattered: agents could access knowledge across the internal knowledge base through the same AI assistant they used during ticket handling, instead of relying solely on manual search or asking colleagues.
AI Assistance Inside Existing Support Workflows
Instead of creating another application for agents to learn, CodeIT integrated AI assistance into the client’s existing management, billing, and communication workflows.
Within the management and billing platform, agents receive AI-generated draft replies and can request summaries of the current ticket state.
A Ticket Assistant inside the team’s existing messenger provides capabilities for:
Context
Current ticket state, next steps, and risk markers.
Knowledge assistance
Natural-language questions answered from internal documentation with sources.
Draft replies
Responses generated from ticket history, internal notes, and relevant knowledge-base content.
Internal-note summaries
Condensed context from information customers do not see.
Handoffs
Generated context written back into the ticket for the next agent.
Fact-checking
Agent drafts checked against internal notes.
Thread summaries
Long ticket-related discussions in Google Chat condensed into key points, helping agents understand the conversation without reading the full thread.
Why it mattered: agents could access ticket context, internal knowledge, and summaries of ongoing discussions within their existing support workflow instead of moving between tools or manually reviewing lengthy conversations.
Private AI Architecture and Production Guardrails
For this client, the core challenge was not simply whether an LLM could generate useful support content. It was whether AI could be used without compromising data boundaries or giving the model authority it should not have.

Self-hosted inference
Embedding and generation models run inside the client’s VPN perimeter. Customer data does not need to reach a third-party AI provider.

Grounded response verification
Internal notes are treated as the source of truth when checking agent drafts. Where information requires human verification, the system uses an explicit placeholder instead of generating a plausible but unverified fact.

Human-controlled AI
The assistant cannot send messages to customers or autonomously change ticket status, close tickets, or assign work. Its write capability is limited to internal notes and was gated behind an allowlist during rollout.

Feedback built into production use
Every generated answer includes helpful/not-helpful feedback tied to that specific response, providing a direct signal for identifying weaknesses from actual usage.
The result is not an autonomous support agent. It is a controlled AI layer designed to help human agents work with better context, grounded knowledge, and explicit verification.
Business Impact
AI assistance became part of day-to-day support operations
Rollout to the full support organization began in September 2026. The implementation introduced several direct operational changes:
One-command ticket context
Agents can retrieve the current state of a ticket without reviewing the entire conversation and internal-note history.
Structured support handoffs
Context summaries can be written back into tickets so the next agent receives a briefing rather than reconstructing the case.
Factual verification before reply
Drafts can be checked against internal notes, helping identify unsupported claims before they reach customers.
Accessible internal knowledge
Agents can ask questions in plain language and receive answers with supporting sources.
Training without customer exposure
New agents can practice through scored simulations before working with live customer tickets.
Measurable AI Response Quality
Combined LLM-as-a-judge evaluation of sampled responses against defined criteria with direct human feedback from agents on whether generated answers were helpful.
Support agents
~45
Regions
3
Operations
24/7
What Made This Solution Different
Private AI by architecture, not as an afterthought
Data isolation shaped the architecture from the start: locally hosted generation and embedding models, retrieval, vector storage, and orchestration operate without dependency on commercial AI APIs.
Domain-Tuned Analysis Instead of Generic AI
Language resources were calibrated specifically for infrastructure support signals such as urgency, threat, temporal pressure, and politeness.
Grounding designed for justified skepticism
Source-attributed retrieval, internal notes as ground truth, explicit unknowns, fact-checking, and human-controlled actions were treated as core system behavior rather than optional safeguards.
AI designed around teams, not individual prompts
Shared ticket context and note transfer reflect the reality of 24/7 support operations, where work moves between people and regions.
Technology stack
- Python
- FastAPI
- Django
- Self-hosted LLM
- PostgreSQL
- ClickHouse
- Redis
- Apache Airflow
- NLP
- Kubernetes
- VPN perimeter
Team
- Python engineer
- PHP engineer
- Frontend engineer
- Delivery manager
- CTO
Explore related services
Custom AI Development
Build AI systems designed around your data, your infrastructure, and the way your teams already work.
Generative AI Development
Deploy generative AI and LLMs on your own infrastructure, with full control over where data goes.
AI Knowledge Assistant
Turn scattered internal documentation into answers your team gets in plain language, with sources attached.
Generative AI & LLM Consulting
Decide what to build, which models to run, and whether to self-host, before development starts.