Building AI apps
AI features need three things from a web framework: a server side that can hold API keys and call a model, a way to show the answer while it's still being generated, and safe rendering of what the model wrote. PyWeb has all three, without a separate frontend.
#Start from the template
pyweb new mychat --template ai-chat
cd mychat
pyweb dev app.pywebThe app runs straight away with a built-in demo model. To use a real one, set environment variables before starting it:
ANTHROPIC_API_KEY=... Anthropic (AI_MODEL defaults to claude-sonnet-5-5)
OPENAI_API_KEY=... AI_MODEL=... OpenAI
OPENAI_BASE_URL=http://localhost:11434/v1 \
AI_MODEL=llama3 Ollama, vLLM, LM Studio or any OpenAI-compatible serverThe template calls the provider's HTTP API with the standard library, so there's nothing else to install. Swap in an official SDK if you prefer: the server function only has to yield text.
#Streaming the answer
A @server function that yields streams each value to the browser as it's produced (see Streaming results). This is the core of the template, slightly shortened:
import os
from pyweb import App, Markdown, RPCError, server
app = App(title="AI chat")
def ask_model(messages):
"""Yield pieces of text from your model provider (see the template for real ones)."""
for word in ("Streaming ", "**works**."):
yield word
@server
def reply(messages: list):
if not messages or messages[-1].get("role") != "user":
raise RPCError("validation_error", "the last message must be from the user")
yield from ask_model(messages[-20:])
@app.page("/")
def Chat():
messages = []
draft = ""
busy = False
stream = None
async def send():
if not draft.strip() or busy:
return
messages.append({"role": "user", "content": draft})
messages.append({"role": "assistant", "content": ""})
draft = ""
busy = True
stream = reply(messages[:-1])
async for piece in stream:
messages[-1]["content"] += piece
busy = False
def stop():
if stream:
stream.cancel()
def on_unmount():
stop()
<main>
for m in messages:
<div class={"bubble " + m["role"]}>
<Markdown text={m["content"]} />
</div>
<form onsubmit={send}>
<input bind={draft} placeholder="Message" />
if busy:
<button type="button" onclick={stop}>Stop</button>
else:
<button>Send</button>
</form>
</main>What this compiles to
// Chat.js
import { markdown as $markdown } from "./markdown.js";
function Chat($s) {
const messages = $signal("messages" in $s ? $s["messages"] : []);
const draft = $signal("draft" in $s ? $s["draft"] : "");
const busy = $signal("busy" in $s ? $s["busy"] : false);
const stream = $signal("stream" in $s ? $s["stream"] : null);
async function send() {
let piece;
if ((!$py.truth($py.m(draft(), "strip")) || $py.truth(busy()))) {
return;
}
$py.mut(messages, [], ($v) => $py.m($v, "append", {"role": "user", "content": draft()}));
$py.mut(messages, [], ($v) => $py.m($v, "append", {"role": "assistant", "content": ""}));
draft("");
busy(true);
stream($rpc.stream("reply", {"messages": $py.slice(messages(), null, (-1), null)}));
for await (piece of $py.aiter(stream())) {
{ const $k1 = [(-1), "content"]; $py.setp(messages, $k1, $py.add($py.getp(messages.peek(), $k1), piece)); }
}
busy(false);
}
function stop() {
if ($py.truth(stream())) {
stream().cancel();
}
}
function on_unmount() {
stop();
}
$onCleanup(on_unmount);
return [$h("main", null, () => [
$list(() => messages(), (m) => [$h("div", {"class": $py.add("bubble ", $py.at(m, "role"))}, () => [$markdown($py.at(m, "content"), null)])]),
$h("form", {"onsubmit": send}, () => [
$h("input", {"$bind": draft, "placeholder": "Message"}),
$when(() => $py.truth(busy()), () => [$h("button", {"type": "button", "onclick": stop}, () => [$t("Stop")])], () => [$h("button", null, () => [$t("Send")])])
])
])];
}
$mount("Chat", Chat);page Chat route=/
browser messages reactive state: .append() in send(); literal initial value
browser draft reactive state: bound to an input (line 53); literal initial value
browser busy reactive state: assigned in send(); literal initial value
browser stream reactive state: assigned in send(); literal initial value
browser send event handler (compiled to JavaScript)
browser stop event handler (compiled to JavaScript)
browser on_unmount event handler (compiled to JavaScript)
rpc POST /__pyweb/rpc/reply (messages: list) -> Any<main>
<form>
<input value="" placeholder="Message"><button>Send</button>
</form></main>- Each piece is shown as it arrives.
messages[-1]["content"] += piecechanges page state, so the last bubble re-renders. - Stop really stops.
stream.cancel()aborts the request. On the server the generator is closed, which closes the connection to the model provider, so you don't pay for tokens nobody reads. Leaving the page does the same throughon_unmount. - Errors reach the page. Raise
RPCError("unavailable", "...")in the generator (the template does this for provider errors) and theasync forraises it; catch it withexcept RPCError as e.
#Showing model output: <Markdown>
Models answer in Markdown. <Markdown text={...} /> (from pyweb) renders it on the server for the first page load and in the browser as the text changes, using the same rules in both places:
- headings, emphasis, inline and fenced code, lists, quotes, tables, links and images;
- an unclosed code fence runs to the end, so half-streamed code blocks look right;
- it is safe on untrusted text: raw HTML is shown as text, and links and images only accept
http(s),mailtoand relative URLs, so a prompt-injected reply can't run script or plant ajavascript:link.
It renders into <div class="markdown">; add classes with class=.
#Keys, cost and abuse
- Keys stay on the server. Read them with
os.environinside server functions. The compiler stops you from sending a variable named like a secret (api_key,token, ...) to the browser. - Treat the conversation as user input. The browser sends the whole history with each call, so the server function should check roles, trim the length (the template keeps the last 20 messages of up to 4,000 characters) and set the model's
max_tokens. - Limit who can call it.
pyweb serverate-limits RPC calls per client (120 a minute by default). For a public app, also require a login (session.require()) and keep per-user budgets in your database.
#Other AI patterns
- Long jobs with progress: yield progress dicts from the server function and show a bar.
- Background work: start it with
pyweb.jobsandpublish()progress to the page over live updates. - npm packages for AI UIs (syntax highlighting, charts of token usage) work in browser code: see npm packages.