Compile C++ to WebAssembly

You have a C++ library and you want it in a web page, without a server. Emscripten compiles C and C++ to WebAssembly, plus a small JavaScript file that loads it. Here’s the whole process, using whisper.cpp (OpenAI’s Whisper speech-to-text model in plain C++) as a real example.

  • C++
  • WebAssembly
  • Tutorial

What you need

  • Linux, macOS or WSL, with git, cmake, make and Python 3
  • About 2 GB of free disk space for the Emscripten SDK
  • A few GB of RAM. Compiling whisper.cpp on a tiny cloud VM can run out of memory

1. Install Emscripten

The Emscripten SDK (emsdk) brings its own Clang/LLVM, Node.js and Python, so it doesn’t touch your system compilers.

git clone https://github.com/emscripten-core/emsdk.git
cd emsdk
./emsdk install latest
./emsdk activate latest
source ./emsdk_env.sh

The last line only sets up the current terminal. Run it again in every new shell, or add it to your ~/.bashrc. Check that it works with emcc --version.

2. Try it on something small

emcc is a drop-in replacement for gcc/clang. Compile a hello world first, so you know the toolchain works before you add a big project:

// hello.cpp
#include <cstdio>
int main() {
    printf("Hello from C++\n");
}
emcc hello.cpp -o hello.html
python3 -m http.server

Open http://localhost:8000/hello.html and the text appears in the page. You get three files: hello.wasm (the compiled code), hello.js (the loader) and hello.html (a test page). Use -o hello.js to skip the HTML when you have your own page.

3. Build a real project: whisper.cpp

For CMake projects you don’t change the build files. You run CMake through emcmake, which swaps in the Emscripten toolchain:

git clone https://github.com/ggml-org/whisper.cpp
cd whisper.cpp
mkdir build-em && cd build-em
emcmake cmake ..
make -j

whisper.cpp detects Emscripten and also builds its browser example. The result is in build-em/bin/: whisper.wasm/ holds the demo page, and libmain.js is the compiled library. By default the .wasm is embedded inside the JavaScript as base64, so you only ship one file. Add -DWHISPER_WASM_SINGLE_FILE=OFF to the cmake line to get a separate libmain.wasm. Base64 adds about a third to the size, and a separate file lets browsers compile it while it downloads.

4. Expose C++ functions to JavaScript

A compiled library is useless until JavaScript can call it. Emscripten’s embind does that. This is a trimmed version of what whisper.cpp does in examples/whisper.wasm/emscripten.cpp:

#include "whisper.h"
#include <emscripten/bind.h>

static whisper_context * ctx = nullptr;

EMSCRIPTEN_BINDINGS(whisper) {
    // Module.init('whisper.bin') -> load a model file
    emscripten::function("init", emscripten::optional_override(
        [](const std::string & path) {
            ctx = whisper_init_from_file_with_params(
                path.c_str(), whisper_context_default_params());
            return ctx != nullptr;
        }));
}

The link flags matter as much as the code. These come from whisper.cpp’s CMakeLists.txt:

--bind                         # turn on embind
-s USE_PTHREADS=1              # real threads, via Web Workers
-s INITIAL_MEMORY=512MB
-s MAXIMUM_MEMORY=2000MB
-s ALLOW_MEMORY_GROWTH=1       # models are big; let the heap grow
-s FORCE_FILESYSTEM=1          # a virtual file system to put the model in
-s EXPORTED_RUNTIME_METHODS="['print', 'printErr', 'ccall', 'cwrap', 'HEAPU8']"

5. Call it from the page

C++ code expects files, but a browser has none. Emscripten gives it an in-memory file system: write the model into it, then call your function with the path. Audio goes in as 16 kHz mono samples in a Float32Array, which is what Whisper expects.

<script>
  var Module = {
    print: (text) => console.log(text),   // C++ printf output ends up here
    printErr: (text) => console.warn(text),
  };
</script>
<script src="libmain.js"></script>
<script>
  async function loadModel(url) {
    const buf = new Uint8Array(await (await fetch(url)).arrayBuffer());
    Module.FS_createDataFile('/', 'whisper.bin', buf, true, true);
    return Module.init('whisper.bin');
  }
  // later: Module.full_default(instance, audio, 'en', 4, false);
</script>

Get a model with the script in the repo: ./models/download-ggml-model.sh base.en saves models/ggml-base.en.bin. The tiny and base models run comfortably in a browser. small is about the limit, because WebAssembly memory is capped here at 2 GB.

6. Serve it with the right headers

This step catches everyone. Threads in WebAssembly need SharedArrayBuffer, and browsers only allow that on pages that are cross-origin isolated. Without these two headers, the page stops at startup with an error about SharedArrayBuffer:

Cross-Origin-Opener-Policy: same-origin
Cross-Origin-Embedder-Policy: require-corp

whisper.cpp ships a small test server that sets them. Run it from the build-em folder, then open http://localhost:8000/whisper.wasm:

python3 ../examples/server.py

On nginx, add them only to the location that serves the app, because require-corp blocks cross-origin images, fonts and scripts that don’t opt in:

location /whisper/ {
    add_header Cross-Origin-Opener-Policy same-origin;
    add_header Cross-Origin-Embedder-Policy require-corp;
}

Also check that .wasm files are sent as application/wasm. Recent nginx versions do this by default; otherwise the browser falls back to a slower way of loading them.

Things that will bite you

  • The output comes through print. whisper.cpp’s example prints the transcript with printf, so you read it from Module.print, not from a return value. For your own code, return strings through embind instead.
  • SIMD is required. whisper.cpp uses 128-bit WebAssembly SIMD for speed. Every current browser supports it, but old ones fail to load the module.
  • No GPU. This all runs on the CPU. Expect about two to three times faster than real time with tiny or base on a modern laptop.
  • Debug native first. Build and test the C++ normally before you compile it to WebAssembly. Stack traces from a .wasm are much less fun to read.

Wrap-up

The recipe is always the same: install emsdk, build with emcc or emcmake, expose functions with embind, put input files into the virtual file system, and serve with the cross-origin headers if you use threads. For speech to text in particular, there’s a second route: Transcriber runs the same Whisper model through Transformers.js and ONNX Runtime, which can also use the GPU through WebGPU.

← Back to Scratchpad