GL Chat
A document chatbot that runs entirely on self-hosted models
A question-answering assistant over a private library of technical documents, built on open-source language models that run on your own hardware.

The challenge
A team with a library of technical material wants to ask it questions in plain language. The obvious route is a hosted AI service, but that means sending internal documents to a third party, which many organisations cannot do.
GL Chat was built in 2023 to answer one question: can a useful document assistant run with nothing leaving the building?
The requirement
- Answers drawn from the organisation’s own documents, not from the model’s general knowledge.
- The passages behind each answer shown to the user.
- No external AI service. Models and data stay on local infrastructure.
- A second mode suited to programming questions.
- A way to turn a learner’s goal and current level into a study plan.
The solution
Documents are loaded as PDFs, split into overlapping passages and indexed. When someone asks a question, the most relevant passages are retrieved and given to the language model with an instruction to answer only from them, and to say so when the answer is not there.
Every response returns the source passages alongside the answer, so the reader can check it.
Users can switch between two models: a general chat model, and a code-focused model for programming questions. A separate endpoint generates a step-by-step learning path from two inputs, what the person wants to learn and what they already know.
The interface is a web chat application with sign-up and login.
Architecture
Ingestion runs as a separate step. PDFs are split into passages, each passage is turned into an embedding by a small sentence-transformer model, and the result is stored in a FAISS index on disk.
At question time, the API embeds the question, retrieves the closest passages from the index and passes them to the model in a fixed prompt. The models are quantised builds of Llama 2 and Code Llama, loaded locally and run on CPU, so the system needs no GPU and makes no calls to an outside service.
A Flask API exposes the question and learning-path endpoints to a React frontend.
Technology
Python, LangChain, Llama 2 and Code Llama (quantised, running locally), sentence-transformers for embeddings, FAISS, Flask, React and Material UI.
Outcome
A working prototype that showed retrieval-grounded answers with visible sources can be delivered on self-hosted open models. The approach it tested, retrieve first and answer only from what was retrieved, is the one we still use for document assistants, with larger models where a client’s data policy allows.

