Transformers.js runs AI models from Hugging Face directly in the browser, with no server, no Python and no API key. The model downloads once, then runs on the visitor’s own device. Here’s how to use it, plus a list of models worth trying.
What you need
- A web page served over HTTP.
python3 -m http.serveris enough; opening the file directly won’t work - A modern browser. Chrome and Edge are fastest, because they support WebGPU
- No build step and no npm. Everything loads from a CDN
1. Your first model in three lines
<script type="module">
import { pipeline } from 'https://cdn.jsdelivr.net/npm/@huggingface/transformers@4.3.0/dist/transformers.min.js';
const classify = await pipeline('sentiment-analysis', 'Xenova/distilbert-base-uncased-finetuned-sst-2-english');
console.log(await classify('I love how fast this page loads!'));
// [{ label: 'POSITIVE', score: 0.9997 }]
</script>
Two details save a lot of confusion. Import dist/transformers.min.js: it’s a complete bundle, while the default transformers.web.js imports other packages by name, which a browser can’t resolve without a bundler. And always pin the version in the URL, so an update can’t break your page.
2. How pipeline() works
Everything goes through one function: pipeline(task, model, options). The task says what kind of job it is ('sentiment-analysis', 'object-detection', 'automatic-speech-recognition'…), and the model is a Hugging Face repository name. You get back an async function you can call as many times as you like.
To find models, browse huggingface.co/models?library=transformers.js. Only models converted to ONNX work, which is why most names start with Xenova/ or onnx-community/. Check each model’s license before you use it in a product; some are for non-commercial use only.
The first call downloads the model files; after that, the browser keeps them in its cache. Small models are a few tens of megabytes, while language models start at a few hundred, so pick the smallest one that does the job.
3. Choose the hardware and the download size
Two options control speed and size. device picks where the model runs: 'webgpu' uses the graphics card, and 'wasm' uses the processor and works everywhere. dtype picks how compressed the weights are. 'q8' and 'q4' are 8-bit and 4-bit versions: much smaller, and usually only a little less accurate.
const device = (await navigator.gpu?.requestAdapter()) ? 'webgpu' : 'wasm';
const classify = await pipeline('sentiment-analysis', 'Xenova/distilbert-base-uncased-finetuned-sst-2-english', {
device,
dtype: 'q8',
});
WebGPU can exist and still fail (old drivers, missing features), so wrap the load in try/catch and fall back to 'wasm'.
4. Run it in a Web Worker
A model blocks whatever thread it runs on. On the main thread, the page freezes until it finishes. Move it into a worker and send messages back and forth:
// worker.js
import { pipeline } from 'https://cdn.jsdelivr.net/npm/@huggingface/transformers@4.3.0/dist/transformers.min.js';
let classify;
self.onmessage = async ({ data }) => {
classify ??= await pipeline('sentiment-analysis', 'Xenova/distilbert-base-uncased-finetuned-sst-2-english', {
progress_callback: (p) => self.postMessage({ type: 'progress', ...p }),
});
self.postMessage({ type: 'result', result: await classify(data.text) });
};
// main page
const worker = new Worker('worker.js', { type: 'module' });
worker.onmessage = ({ data }) => console.log(data);
worker.postMessage({ text: 'This is great' });
progress_callback reports every model file separately, with file, loaded and total. Add them up yourself to show one download bar.
5. What goes in and what comes out
- Text goes in as a string, or an array of strings to process several at once.
- Images can be a URL, a blob URL from a file input, or a canvas. Image results such as cut-outs and depth maps come back as a
RawImage, withtoCanvas()andtoBlob()to show or save them. - Audio goes in as a
Float32Arrayat the sample rate the model expects (16 kHz for speech models). Generated speech comes back withtoBlob(), ready for an<audio>element.
6. Models worth trying
Each is a single pipeline() call. I ran all of them in a browser with Transformers.js 4.3, except the chat model, which was too much for my small test server.
Text
- Sorting text into your own categories zero-shot-classification ·
Xenova/mobilebert-uncased-mnli - No training needed: you give it the labels when you call it. Good for routing support emails or tagging notes.
- Finding names, places and companies ner ·
Xenova/bert-base-NER - Marks people, organizations and locations in a sentence.
- Search by meaning feature-extraction ·
Xenova/all-MiniLM-L6-v2 - Turns text into 384 numbers, so that similar meanings end up close together. It’s the basis for semantic search and “related articles” without a server.
- A small chat model text-generation ·
HuggingFaceTB/SmolLM2-360M-Instruct - A real instruction-following language model that runs offline. Don’t expect much knowledge, but it’s fine for rewording and short answers.
const classify = await pipeline('zero-shot-classification', 'Xenova/mobilebert-uncased-mnli');
await classify('My invoice was charged twice this month', ['billing', 'technical issue', 'sales']);
// { labels: ['billing', 'technical issue', 'sales'], scores: [0.57, 0.40, 0.03] }
const embed = await pipeline('feature-extraction', 'Xenova/all-MiniLM-L6-v2');
const out = await embed(['How do I reset my password?', 'I forgot my login'], { pooling: 'mean', normalize: true });
const [a, b] = out.tolist();
a.reduce((sum, x, i) => sum + x * b[i], 0); // 0.61: similar meaning, no shared words
Images
- Removing the background background-removal ·
Xenova/modnet - Cuts the person out of a portrait and returns an image with a transparent background.
- Finding objects in a photo object-detection ·
Xenova/detr-resnet-50 - Returns a label, a confidence score and a box for everything it recognizes.
- Estimating depth depth-estimation ·
onnx-community/depth-anything-v2-small - Guesses how far away every pixel is from a single ordinary photo. Fun for 3D effects.
- Describing a photo image-to-text ·
Xenova/vit-gpt2-image-captioning - Writes a short caption. Handy for suggesting alt text.
const removeBg = await pipeline('background-removal', 'Xenova/modnet');
const cutout = await removeBg('portrait.jpg');
document.body.append(cutout.toCanvas());
const detect = await pipeline('object-detection', 'Xenova/detr-resnet-50');
await detect('cats.jpg', { threshold: 0.9 });
// [{ label: 'cat', score: 0.99, box: { xmin, ymin, xmax, ymax } }, ...]
const caption = await pipeline('image-to-text', 'Xenova/vit-gpt2-image-captioning');
await caption('cats.jpg');
// [{ generated_text: 'two cats laying on a bed with a remote' }]
Audio
- Speech to text automatic-speech-recognition ·
onnx-community/whisper-base - OpenAI’s Whisper, in many languages. See the example below.
- Text to speech text-to-speech ·
Xenova/mms-tts-eng - A small English voice from Meta’s MMS project. Robotic, but tiny and fast.
const speak = await pipeline('text-to-speech', 'Xenova/mms-tts-eng');
const speech = await speak('Hello from your browser.');
new Audio(URL.createObjectURL(await speech.toBlob())).play();
Example: how Transcriber is built
My Transcriber is all of the above put together for speech to text: Whisper in a Web Worker, WebGPU with a fallback to WebAssembly, and a download progress bar. The only speech-specific parts are preparing the audio and handling long recordings:
// The browser decodes almost any file; an offline context resamples it to 16 kHz mono.
const decoded = await new AudioContext().decodeAudioData(await file.arrayBuffer());
const offline = new OfflineAudioContext(1, Math.ceil(decoded.duration * 16000), 16000);
const src = offline.createBufferSource();
src.buffer = decoded;
src.connect(offline.destination);
src.start();
const audio = (await offline.startRendering()).getChannelData(0);
const transcribe = await pipeline('automatic-speech-recognition', 'onnx-community/whisper-base');
const result = await transcribe(audio, {
language: 'en',
chunk_length_s: 30, // Whisper sees 30 seconds at a time
stride_length_s: 5, // overlap the windows so words at the edges aren't cut
return_timestamps: true, // result.chunks becomes subtitle cues
});
Things that will bite you
- Not every task exists. Version 4 has no translation or summarization pipelines. Check the supported tasks before you plan a feature around one.
- Defaults can be surprising. Whisper, for example, assumes English when you don’t pass a language, instead of detecting it. Transcriber adds its own detection step.
- WebAssembly runs on one core. Multi-threading needs the page to be cross-origin isolated, with the same headers as in the C++ to WebAssembly article. Without them, CPU-only devices are slow with bigger models.
- Cached scripts hide your changes. If your server caches
.jsfiles, browsers can keep running an old worker. Put a version in the URL (worker.js?v=2) and bump it on every change. - Phones have less memory. A model that runs fine on a laptop can crash a mobile tab. Offer a smaller model, or say up front how big the download is.
Wrap-up
Import the bundle from a CDN, pick a task and a model, run it in a Web Worker, use WebGPU when it’s there, and choose the smallest model that works. After that, swapping in a different model is usually a one-line change.