Home Our Work Self-Hosted AI Support Layer

Self-Hosted AI Layer for Customer Support

Location

USA, Czech Republic

Partnership period

2022 – Ongoing

Team size

5

Project information

Overview

A global infrastructure hosting provider operates data centers across North America, Europe, and Asia, serving customers from individual developers to enterprises running production workloads.

Its 24/7 support organization includes approximately 45 agents across three regions. The engagement began during a period of rapid growth, as ticket volume was increasing faster than headcount and the company was transforming its support operations across process, staffing, and tooling.

Self-hosted AI layer for customer support

Challenge

AI training simulator for support agents

Scaling support without compromising data control or response quality

As support operations expanded across regions, a significant part of the workload came before agents could start solving the actual customer issue. They had to reconstruct ticket history, search across fragmented internal knowledge, and coordinate work across shifts.

At the same time, tickets contained customer data and infrastructure details, which meant commercial LLM APIs did not meet the client’s requirements.

The key challenges were:

Complex cross-shift handoffs

Follow-the-sun support meant multiple agents could work on the same ticket, requiring each new agent to reconstruct its context.

Ticket volume outpacing headcount

Growing support demand increased the operational impact of time lost during repeated handoffs.

Fragmented internal knowledge

Answers were distributed across the knowledge base, billing platform, and internal notes, with limited search across them.

Inconsistent and error-prone responses

Response structure and completeness varied, while unsupported claims could lead to rework and escalation.

No safe path to public AI

Customer and infrastructure data could not be sent to commercial LLM APIs.

The client needed to reduce the time agents spent getting into a ticket, make internal knowledge answerable, and keep customer data entirely within the company’s own network.

Our Approach

Build AI around the support operation—not the other way around

Working closely with the client’s CTO, CodeIT designed a self-hosted AI layer around the existing support environment rather than introducing a standalone AI tool.

The architecture followed four principles from the start:

  • Keep AI private. Generation and embeddings run inside the client’s VPN perimeter, with no customer data sent to third-party AI providers.
  • Ground AI in company knowledge. Internal documentation and agent notes provide the evidence behind generated answers and fact-checking.
  • Keep people in control. The assistant does not autonomously send customer responses or change, close, or assign tickets.
  • Fit existing workflows. AI capabilities were embedded into the systems and team workflows agents already used rather than requiring them to move to a separate AI application.

This resulted in a five-component AI support layer running entirely within the client’s network perimeter.

Self-Hosted AI Support Layer

The Solution

Ticket Analysis & Context Generation

CodeIT built a multi-stage enrichment pipeline that converts raw ticket text into structured support context.

The pipeline identifies signals such as urgency, politeness, and escalation pressure; extracts technical artifacts including IP addresses, ASNs, error codes, and server identifiers; and uses locally hosted models to generate structured summaries covering the customer’s request, work already performed, current blocker, and escalation risk.

Language resources were tuned specifically for hosting and infrastructure support rather than relying only on generic sentiment analysis.

Why it mattered: agents could work from structured ticket context instead of rebuilding the situation entirely from conversation history.

AI training simulator for support agents

Knowledge Base Access Through the AI Assistant

The AI Assistant gives support agents conversational access to the client’s internal knowledge base, including years of runbooks, procedures, hardware guidance, and network documentation.

Agents can ask questions in natural language and receive answers based on relevant internal content, together with links to the supporting source articles. Behind the assistant, a RAG pipeline retrieves relevant information from the indexed knowledge base and grounds generated answers in company documentation.

Why it mattered: agents could access knowledge across the internal knowledge base through the same AI assistant they used during ticket handling, instead of relying solely on manual search or asking colleagues.

AI Assistance Inside Existing Support Workflows

Instead of creating another application for agents to learn, CodeIT integrated AI assistance into the client’s existing management, billing, and communication workflows.

Within the management and billing platform, agents receive AI-generated draft replies and can request summaries of the current ticket state.

A Ticket Assistant inside the team’s existing messenger provides capabilities for:

Context

Current ticket state, next steps, and risk markers.

Knowledge assistance

Natural-language questions answered from internal documentation with sources.

Draft replies

Responses generated from ticket history, internal notes, and relevant knowledge-base content.

Internal-note summaries

Condensed context from information customers do not see.

Handoffs

Generated context written back into the ticket for the next agent.

Fact-checking

Agent drafts checked against internal notes.

Thread summaries

Long ticket-related discussions in Google Chat condensed into key points, helping agents understand the conversation without reading the full thread.

Why it mattered: agents could access ticket context, internal knowledge, and summaries of ongoing discussions within their existing support workflow instead of moving between tools or manually reviewing lengthy conversations.

Private AI Architecture and Production Guardrails

For this client, the core challenge was not simply whether an LLM could generate useful support content. It was whether AI could be used without compromising data boundaries or giving the model authority it should not have.

house

Self-hosted inference

Embedding and generation models run inside the client’s VPN perimeter. Customer data does not need to reach a third-party AI provider.

code-computer

Grounded response verification

Internal notes are treated as the source of truth when checking agent drafts. Where information requires human verification, the system uses an explicit placeholder instead of generating a plausible but unverified fact.

person-add

Human-controlled AI

The assistant cannot send messages to customers or autonomously change ticket status, close tickets, or assign work. Its write capability is limited to internal notes and was gated behind an allowlist during rollout.

revote (1)

Feedback built into production use

Every generated answer includes helpful/not-helpful feedback tied to that specific response, providing a direct signal for identifying weaknesses from actual usage.

The result is not an autonomous support agent. It is a controlled AI layer designed to help human agents work with better context, grounded knowledge, and explicit verification.

Business Impact

AI assistance became part of day-to-day support operations

Rollout to the full support organization began in September 2026. The implementation introduced several direct operational changes:

One-command ticket context

Agents can retrieve the current state of a ticket without reviewing the entire conversation and internal-note history.

Structured support handoffs

Context summaries can be written back into tickets so the next agent receives a briefing rather than reconstructing the case.

Factual verification before reply

Drafts can be checked against internal notes, helping identify unsupported claims before they reach customers.

Accessible internal knowledge

Agents can ask questions in plain language and receive answers with supporting sources.

Training without customer exposure

New agents can practice through scored simulations before working with live customer tickets.

Measurable AI Response Quality

Combined LLM-as-a-judge evaluation of sampled responses against defined criteria with direct human feedback from agents on whether generated answers were helpful.

Support agents

~45

Regions

3

Operations

24/7

What Made This Solution Different

Private AI by architecture, not as an afterthought

Data isolation shaped the architecture from the start: locally hosted generation and embedding models, retrieval, vector storage, and orchestration operate without dependency on commercial AI APIs.

Domain-Tuned Analysis Instead of Generic AI

Language resources were calibrated specifically for infrastructure support signals such as urgency, threat, temporal pressure, and politeness.

Grounding designed for justified skepticism

Source-attributed retrieval, internal notes as ground truth, explicit unknowns, fact-checking, and human-controlled actions were treated as core system behavior rather than optional safeguards.

AI designed around teams, not individual prompts

Shared ticket context and note transfer reflect the reality of 24/7 support operations, where work moves between people and regions.

Technology stack

  • Python
  • FastAPI
  • Django
  • Self-hosted LLM
  • PostgreSQL
  • ClickHouse
  • Redis
  • Apache Airflow
  • NLP
  • Kubernetes
  • VPN perimeter

Team

  • Python engineer
  • PHP engineer
  • Frontend engineer
  • Delivery manager
  • CTO

Explore related services

Business First
Code Next
Let’s talk

    By clicking the “Send” button I confirm, that I have read and agree to the Privacy Policy.