Choose where intelligence runs.
Use the free 135M hosted preview and chat API, or download a 0.5B or 1.5B model to your WebGPU browser. The larger models run on your device.
Try a request
No key. No wallet.Your response appears here. Choose an example or write your own prompt to begin.
Shared preview capacity · 6 requests/minute per instance/IP · No prompts or outputs are saved by Toll. Requests run on the server.
One familiar interface.
Use the chat-completions request format with your existing tooling. A small API with live streaming with clear limits and no API key.
- POST · Chat completions
- /v1/chat/completions
- GET · Model catalog
- /v1/models
- Model identifier
- toll-small-135m
A bigger model, right here.
Download once, run on your device. No API key or paid inference account. Prompts stay in this tab; model files download from their public repository.
Larger downloads and GPU memory are required. Performance depends on the device. These models run in the browser, not through the hosted chat API; no paid fallback is used.