RAG backend and model orchestration for CMU GPT
  • Python 98.7%
  • Nix 1.3%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-09-26 18:22:09 +00:00
.forgejo docs: add contributing guide and rewrite readme 2026-09-18 17:43:13 -04:00
evals style: follow PEP 257 docstrings and PEP 8 naming 2026-09-18 17:00:12 -04:00
src/cmugpt style: follow PEP 257 docstrings and PEP 8 naming 2026-09-18 17:00:12 -04:00
tests/unit style: follow PEP 257 docstrings and PEP 8 naming 2026-09-18 17:00:12 -04:00
.editorconfig Initial commit 2025-09-21 23:39:21 -04:00
.env.example chore: use port 5055 for local runs 2026-09-18 00:57:04 -04:00
.envrc feat: nix 2026-05-30 15:30:06 -04:00
.gitignore docs: add contributing guide and rewrite readme 2026-09-18 17:43:13 -04:00
.python-version Initial commit 2025-09-21 23:39:21 -04:00
CONTRIBUTING.md docs: add contributing guide and rewrite readme 2026-09-18 17:43:13 -04:00
devenv.lock fix: devenv location 2026-08-02 17:33:02 -04:00
devenv.nix chore: use port 5055 for local runs 2026-09-18 00:57:04 -04:00
devenv.yaml chore: update kennel deployment 2026-08-12 21:31:00 -07:00
flake.lock fix: flake and build 2026-08-02 19:20:22 -04:00
flake.nix chore: remove the Procfile and duplicate flake output 2026-09-18 01:01:08 -04:00
LICENSE Initial commit 2025-09-21 23:39:21 -04:00
pyproject.toml style: follow PEP 257 docstrings and PEP 8 naming 2026-09-18 17:00:12 -04:00
README.md docs: setting up devenv 2026-09-26 13:44:36 -04:00
secretspec.toml fix: turn on the production startup check 2026-09-18 00:55:48 -04:00
uv.lock build: upgrade anyio to 4.14.2 for security advisories 2026-09-18 18:09:57 -04:00

Bark Agent

Bark Agent is the backend service for Bark, the campus assistant for Carnegie Mellon University built by ScottyLabs. It receives chat messages from the Bark web application and answers them with a language model and a set of campus data tools. Each answer is checked for safety and accuracy before it is returned, and the service keeps long-term memory for each user.

Overview

Bark consists of two services.

  • The Surface (cmugpt-surface) is the web application and its server. It authenticates users, stores chats, and forwards each message to this service.
  • The Agent (this repository) processes each message. It selects the campus tools the question requires, runs a LangGraph agent against models served through OpenRouter, validates the result, and streams the answer back to the Surface.

Campus data comes from the CMU MCP server (mcp-server), which publishes tools for maps, courses, dining, and the student guide over the Model Context Protocol. OpenAI provides the embeddings used for memory search and the moderation endpoint. Per-user memory is stored in PostgreSQL with the pgvector extension.

flowchart TB
    browser["Browser"]
    surface["Surface<br/>web app and API server"]
    agent["Bark Agent<br/>(this repository)"]
    openrouter["OpenRouter<br/>language models"]
    mcp["CMU MCP server<br/>campus data tools"]
    openai["OpenAI<br/>embeddings, moderation"]
    postgres["PostgreSQL<br/>user memory (pgvector)"]

    browser -->|chat| surface
    surface -->|"POST /agent/respond/stream"| agent
    agent --> openrouter
    agent --> mcp
    agent --> openai
    agent --> postgres

    style agent stroke-width:3px

Request lifecycle

A request passes through five stages.

  1. Validation. The Surface posts the message, the prior turns of the chat, and a hashed user identifier. Before any model call, the service enforces request size limits, checks the user's daily token budget, and screens the message with OpenAI's moderation endpoint.
  2. Planning. planning.py determines what the turn requires. It decides which tool groups to bind, whether the remember and forget tools are needed, and whether memory recall should run. Conversational messages bind no data tools. The map tool is bound on every turn unless the user has disabled maps, because the model decides whether a map belongs in the answer.
  3. Execution. graph.py runs a LangGraph graph. It recalls relevant facts about the user, invokes the model, executes any tool calls, and repeats until the model produces a final answer. Tool output is wrapped as untrusted data so that it cannot inject instructions.
  4. Verification. guards.py and the maps/ package check the finished answer without a model. The model's map selection is validated against the building catalog, and incorrect claims that a lookup failed are repaired. Secrets and system prompt text are removed, and tool usage is disclosed accurately.
  5. Delivery. The answer is streamed to the Surface as Server-Sent Events. Once the answer is complete, a background task extracts durable facts about the user from the exchange and stores them for future turns.

Memory holds only the facts extracted from conversations and the facts a user explicitly asks Bark to remember. Raw chat turns are never stored. A user can view and delete their facts through the Surface, and each user's memory is kept separate by identifier.

Project structure

cmugpt-agent/
│
├── src/cmugpt/
│   │
│   ├── api/                    # HTTP layer
│   │   ├── server.py           #   FastAPI app: lifespan, CORS, error envelope, routers
│   │   ├── deps.py             #   Bearer-token and body-size checks shared by routes
│   │   └── routes/
│   │       ├── agent.py        #   /agent/respond, /agent/respond/stream, /agent/title
│   │       ├── memory.py       #   /memory/*
│   │       └── health.py       #   /api/health
│   │
│   ├── settings.py             # All environment variables
│   ├── schema.py               # Request and response models
│   ├── planning.py             # Per-turn tool and memory selection
│   ├── graph.py                # LangGraph control flow
│   ├── prompts.py              # System prompt construction
│   ├── llm.py                  # OpenRouter model client factory
│   ├── mcp_tools.py            # MCP tool discovery and group filtering
│   ├── guards.py               # Deterministic output checks
│   ├── moderation.py           # OpenAI moderation for input and output
│   ├── token_limits.py         # Per-user daily token budget (SQLite)
│   ├── title.py                # Chat title generation
│   │
│   ├── memory/                 # Per-user long-term memory
│   │   ├── store.py            #   LangGraph store: Postgres with pgvector, or in-memory
│   │   ├── facts.py            #   Recall, save, forget
│   │   ├── tools.py            #   The remember and forget tools exposed to the model
│   │   ├── extraction.py       #   Background fact extraction
│   │   └── manage.py           #   List, delete, clear
│   │
│   └── maps/                   # Campus map support
│       ├── buildings.py        #   Building catalog and aliases
│       ├── buildings.json      #   Catalog data
│       ├── tool.py             #   The maps_show_map tool
│       └── inference.py        #   Map validation and URL construction
│
├── tests/unit/                 # Offline tests, run by CI
├── evals/                      # Live evaluations, run manually
│
├── pyproject.toml              # Dependencies, entry point, tool configuration
├── devenv.nix                  # Local environment and Kennel settings
├── flake.nix                   # Nix build
├── secretspec.toml             # Secret declarations
└── .env.example                # Environment variable template

Requirements

You will need the following, or Nix and devenv in their place (see Installation with devenv):

Installation

The steps below set up local development with the full feature set: persistent memory, semantic memory search, and moderation. This configuration is recommended, since the service then behaves as it does in production. The service also starts without OPENAI_API_KEY or DATABASE_URL, with the reduced behavior described under Configuration.

  1. Clone the repository and install its dependencies.

    git clone https://git.cmu.dev/ScottyLabs/cmugpt-agent.git
    cd cmugpt-agent
    uv sync
    
  2. Create the memory database. Install PostgreSQL and the pgvector extension, then create an empty database named cmugpt_agent. The service creates the pgvector extension, its schema, and its tables on first start. If you prefer not to install PostgreSQL by hand, the devenv shell defined in devenv.nix provides a database with pgvector already set up (see Installation with devenv).

  3. Create the environment file and add the API keys.

    cp .env.example .env
    

    Then set the two keys in .env:

    OPENROUTER_API_KEY=<your OpenRouter key>
    OPENAI_API_KEY=<your OpenAI key>
    

    MCP_SERVER_URL and DATABASE_URL are prefilled. They point at the production MCP server and at the cmugpt_agent database on the local default socket.

Installation with devenv

Members of the slai team can use the devenv shell in place of the Installation steps. It provides Python, uv, and PostgreSQL with pgvector, and it loads the team's shared development keys from OpenBao, so no personal API keys or .env are needed.

Before starting, complete the ScottyLabs setup on docs.scottylabs.org:

  • Forgejo Setup creates a git.cmu.dev account and SSH key.
  • Contributing explains how to join a team through governance. Join slai (data/teams/slai.toml), since team membership grants access to the secrets.
  • Credentials describes how ScottyLabs stores secrets in OpenBao. Kennel's Secrets guide covers the local login and secretspec profiles used below.

Then install the tools:

  1. Nix, with the Determinate Systems installer on macOS, Linux, or WSL:

    curl -fsSL https://install.determinate.systems/nix | sh -s -- install
    
  2. devenv, from a new terminal:

    nix profile install nixpkgs#devenv
    
  3. Optionally, direnv with its shell hook, so the environment loads when you enter the repository.

The shell resolves secrets as it starts and fails without an OpenBao token, so log in once per machine first:

nix run git+https://git.cmu.dev/ScottyLabs/kennel#login

Then clone the repository and start PostgreSQL and the service together:

git clone ssh://forgejo@git.cmu.dev/ScottyLabs/cmugpt-agent.git
cd cmugpt-agent
devenv up

The token renews on each shell entry and expires after 90 days without use. If the shell fails with an OpenBao or permission error, run the login command again. If it still fails, confirm that your git.cmu.dev username is listed in data/teams/slai.toml. secretspec check -P dev reports which secrets resolve without printing their values.

Configuration

All configuration is read from environment variables by settings.py, and .env.example documents each variable with its default.

Variable Purpose
OPENROUTER_API_KEY Chat, memory extraction, and chat titles
OPENAI_API_KEY Embeddings for memory search and the moderation endpoint. Unset, recall orders facts by recency and moderation is skipped
DATABASE_URL PostgreSQL connection string. Unset, memory lives in an in-memory store that is cleared on restart
MCP_SERVER_URL Base URL of the CMU MCP server, including the /mcp path
AGENT_SHARED_SECRET Bearer token the Surface presents on every request. Unset, requests are unauthenticated. Set in production
AGENT_ENV production makes startup fail without DATABASE_URL and an AGENT_SHARED_SECRET of at least 32 characters. The prod profile of secretspec.toml sets it
ALLOWED_ORIGINS Comma-separated browser origins for CORS. Default https://cmugpt.com
PORT Listening port. Default 5055
TITLE_MODEL Model for chat titles. Default qwen/qwen3.7-flash
MEMORY_EXTRACTION_MODEL Model for background fact extraction. Default qwen/qwen3.7-flash
TOKEN_USAGE_DB SQLite file for the daily token budget. Default /tmp/cmugpt_token_usage.sqlite3

Production does not read .env. Kennel injects DATABASE_URL for its managed PostgreSQL instance, the API keys and MCP_SERVER_URL are resolved from OpenBao, and AGENT_SHARED_SECRET is set so that only the Surface can call the service. See Deployment.

Running the service

Start the service with:

uv run cmugpt-agent

It listens on port 5055. To confirm that it is healthy, request the health route:

curl -s localhost:5055/api/health
{"status":"ok","memory":{"backend":"postgres","initialized":true,"semantic_search":true,"embedding_model":"text-embedding-3-large","ready":true}}

backend reports in-memory when DATABASE_URL is unset, and semantic_search is false when OPENAI_API_KEY is unset. When the memory store cannot be queried, the endpoint returns HTTP 503 with "status": "degraded".

API

When AGENT_SHARED_SECRET is set, every route except /api/health requires the header Authorization: Bearer <secret>.

Route Description
POST /agent/respond Returns the complete answer as a JSON object
POST /agent/respond/stream Returns the answer as Server-Sent Events
POST /agent/title Generates a short title from a chat's first message
GET /memory/{user_id} Lists a user's stored facts, with search and paging
DELETE /memory/{user_id}/items/{kind}/{item_id} Deletes one fact
DELETE /memory/{user_id} Deletes all facts for a user
GET /api/health Service status and active memory backend

For example:

curl -s localhost:5055/agent/respond \
  -H 'content-type: application/json' \
  -d '{"query": "What is open for lunch near Gates?", "user_id": "example"}'

Request fields for /agent/respond and /agent/respond/stream:

Field Description
query The user's message. Required. At most 8,000 characters
user_id Identifier for memory and the token budget. The Surface sends a hash of the authenticated user
message_history Prior turns as {"role", "content"} objects. The last 40 are used
model OpenRouter model identifier. Default openai/gpt-5.6-luna
disabled_tools Tool groups the user has switched off: maps, courses, eats, guide

The streaming endpoint emits several event types. status events are sent while tools run, and delta events carry text as it is generated. A map event is sent when a campus map accompanies the answer, and a memory event when a fact is saved or removed. A final done event contains the complete response object, and an error event terminates a failed turn.

Each user is limited to one million tokens per day. Requests beyond that limit receive HTTP 429.

Testing

DATABASE_URL="" uv run pytest    # offline unit tests, as run by CI
uv run pytest evals              # live evaluations

The unit tests in tests/unit/ run with the model replaced by a stub and an in-memory store. They are deterministic and fail only when the code is incorrect.

The evaluations in evals/ send real questions to the configured model and MCP server and check the behavior of the answers: tool usage, refusal of prompt injection, and absence of fabricated details. They require OPENROUTER_API_KEY and MCP_SERVER_URL and incur API costs. When either key is absent they are skipped, and they are not part of the default pytest run.

Development

The code is formatted and linted with ruff and type-checked with ty:

uv run ruff format    # format
uv run ruff check     # lint. The project configuration applies fixes.
uv run ty check       # type check

Deployment

Production runs on Kennel, the ScottyLabs deployment platform. Kennel builds the agent package defined in flake.nix and runs its cmugpt-agent entry point as a systemd unit. PORT, DATABASE_URL, and the secrets from the prod profile of secretspec.toml are injected as environment variables. That profile sets AGENT_ENV=production, so the service refuses to start without DATABASE_URL and an AGENT_SHARED_SECRET of at least 32 characters. Pushes to main on git.cmu.dev trigger a deployment, and each pull request receives a preview deployment.

Production secrets are stored in OpenBao and managed with secretspec.

Contributing

CONTRIBUTING.md describes the workflow, the code style, the commit conventions, and the pull request process.

  • cmugpt-surface: the web application and its server.
  • mcp-server: the CMU MCP server that publishes the campus data tools.
  • kennel: the deployment platform.

License

Apache License 2.0. See LICENSE.