← Personal projects Applied AI Proof of concept · 2026
One assistant, three roles, and a tool layer that says no
A finance assistant where employees, clients and the owner all talk to the same chatbot and each sees only what their role allows. A small model routes each message to one of three agents, and authorization lives in the tool code, never in the prompt.
- 3
- Specialist agents
- 3
- Roles
- pgvector, keyword fallback
- Retrieval
- Swappable — hosted or local
- Model
Architecture
- 01Chat UIstreaming
- 02Auth contextsigned session → role
- 03Routermini-model, 3 intents
- 04Agentfinance · documents · actions
- 05Toolsrole-scoped, audited
- 06PostgresSupabase + pgvector
The problem
A small company’s finance questions come from three kinds of people. An employee wants to know why their payslip is lower this month. A client wants a copy of an overdue invoice. The owner wants to know the cash position and who to chase. One assistant should handle all three — and an employee must never be able to talk it into showing someone else’s salary.
That last sentence is the whole design problem. Building a chatbot that answers finance questions is straightforward. Building one that can’t be argued out of its permissions is not.
Authorization is not a prompt
The tempting approach is to put the rules in the system prompt: you are talking to an employee, only discuss their own records. The prompt in this project does say that, because it makes the model’s refusals read naturally. But nothing depends on it.
The session is resolved to a user and a role on the server before any model is called. The
tools each agent receives are built for that role: an employee’s get_payslips tool is
constructed with their employee ID already bound, and there is no argument the model can pass
to ask for anyone else’s. The owner’s tool set contains company-wide queries; the employee’s
doesn’t contain them at all.
So a successful prompt injection gets the model to want to fetch another person’s data, and then discover it has no tool that can.
Why route instead of one big agent
Each message first goes to a small, cheap model that classifies it into one of three intents, and the conversation is handed to the matching agent:
- Structured finance — reads numbers from Postgres: payslips, invoices, balances, cash.
- Document knowledge — retrieves policy text by vector search: expense rules, payment terms.
- Actions — the only agent that writes: submit a claim, request an advance, queue a reminder.
Three reasons for the split, in the order they mattered:
- Reliability. Each agent gets a short prompt and only the three to six tools it needs. Tool selection is noticeably better with a short menu, especially on small or local models.
- Safety. The agent that can write is only invoked when the user asked for an action. A question about policy can’t accidentally submit a claim.
- Cost. Classification is about fifty tokens on a mini model.
If classification fails — some local models don’t support structured output — a keyword heuristic takes over, and the default is the read-only finance agent.
Retrieval with a fallback
Policy documents are embedded into pgvector and retrieved by similarity. Until the embedding step has been run, the same tool falls back to keyword search, so the application works on a fresh database and gets better when the vectors exist.
Uploaded files — a receipt, an invoice — are stored immediately, filed into the right folder by a tool call, and receipt images are read by the model to pre-fill an expense claim.
What I’d do differently
An evaluation set for the router. Three intents and a heuristic fallback is easy to test: a few hundred labelled messages would give routing accuracy, and show which phrasings land in the wrong agent.
Adversarial tests for the tool layer. The claim that permissions hold under prompt injection should be a test suite, not a paragraph — a set of hostile prompts per role, asserting on which rows the tools returned.
Row-level security in the database. The tools enforce scope in application code. Postgres can enforce the same rule underneath, so that a bug in one tool is not a data leak.