Toll / InferenceINTELLIGENCE FOR YOUR NEXT REQUEST

Choose where intelligence runs.

Use the free 135M hosted preview and chat API, or download a 0.5B or 1.5B model to your WebGPU browser. The larger models run on your device.

Try a request

No key. No wallet.
MODEL RESPONSE

Your response appears here. Choose an example or write your own prompt to begin.

Shared preview capacity · 6 requests/minute per instance/IP · No prompts or outputs are saved by Toll. Requests run on the server.

BUILT FOR INTEGRATION

One familiar interface.

Use the chat-completions request format with your existing tooling. A small API with live streaming with clear limits and no API key.

POST · Chat completions
/v1/chat/completions
GET · Model catalog
/v1/models
Model identifier
toll-small-135m
MORE MODELS. YOUR COMPUTE.

A bigger model, right here.

Download once, run on your device. No API key or paid inference account. Prompts stay in this tab; model files download from their public repository.

WebGPU required

Larger downloads and GPU memory are required. Performance depends on the device. These models run in the browser, not through the hosted chat API; no paid fallback is used.

YOUR ACCESS TO ARC

Your wallet. Your call.

Choose a wallet to get started. Connecting shares your address; it does not approve a payment.

Choose your walletBrowser connection

On mobile, open Toll inside your wallet’s browser. No seed phrase or private key is ever requested.

Arc mainnet · USDC for gas