Local inference, written in daslang.

dasLLAMA is an LLM + speech-to-text inference engine written entirely in daslang and JIT-compiled. It loads stock GGUF files and runs them on cpu, metal, or vulkan. Kernels are auto-tuned per CPU class; the winning variants ship with the engine as class profiles, so a covered box is tuned by its first compile.

12
model architectures, stock GGUF
3
backends — cpu · metal · vulkan
—
measurements on the ladder
01dasllama-server, the OpenAI-compatible server, as a standalone download: macOS, Apple silicon · Windows x64 · Linux x86_64 · Linux arm64. Unpack, start watchdog (the app on a Mac), open 127.0.0.1:8080 and pick a model from the catalog. The control page's benchmark button measures your box (pp512 / tg128 on the served model); dasllama-bench.exe beside the server (inside the Mac app, without the suffix) runs the same rows from a shell, and with --ref pointed at a llama-bench binary the comparison too (how ↗). The archives are not code-signed: macOS asks once under Privacy & Security, Windows SmartScreen once. README ↗
02daslang on github ↗ — the language and SDK this runs on; pre-built binaries at daslang.io/downloads. From the unpacked SDK, bin/daslang -jit utils/dasllama-server/main.das runs the same server through the JIT, which tunes every kernel on your own box instead of picking from the shipped class profiles
03dasLLAMA on github ↗ — the engine, in-tree at modules/dasLLAMA
2026-09-16examples

The storyteller, storywish and parrot now run in every browser - Safari and iOS included.latest

▾

The three examples were compiled to wasm64, which WebKit does not implement, so Safari and every iPhone and iPad browser got a note instead of a story. The engine still cross-compiles as wasm64 inside - that is what keeps its layouts honest - but the link now lowers the memory to 32-bit, and the one SIMD extension WebKit lacks left the build. Try it.

2026-09-11engine

The Qwen family runs on Vulkan.

▾

Every Qwen text model in the zoo, 0.5B to 48 GB, dense and mixture-of-experts, runs on the Vulkan tier. Qwen3.8-27B UD-IQ4_XS runs whole on an RTX 5060 Ti 16 GB under Windows at 865.4 / 23.5 tok/s (pp512 / tg128, measured 2026-09-07), Qwen3.6-35B-A3B UD-IQ2_XXS at 2995.8 / 107.7 there (2026-09-09) and at 5156 / 146.7 (pp512 / tg32, 2026-09-11) on an RTX 5080 under Linux. We have a story to tell about how we got there: Vulkan and the Qwen family fortune.

2026-09-10examples

Storywish - type the words, a story model trained to take requests writes the tale, in the browser.

▾

A second dasLLAMA example on the examples page. The published TinyStories-Instruct models are GPT-Neo, which no GGUF engine runs, so we trained our own: tinystories-instruct-27M, a 27M-parameter llama on the TinyStoriesInstruct corpus with a 4K vocabulary, 15 minutes on one H100. Asked for three words over five word triples, 24 sampled stories each, it puts all three in 61 of the 120 stories; the official 33M puts them in 60. It ships as a 32 MB .dlim beside a 65 MB Kyutai Pocket TTS file with one voice inside and no phoneme packs; you type the words, Enter tells the story, Tab asks for dialogue. The model, the training recipe and the hit-rate numbers are on Hugging Face. The model sets of both examples are now minted by the deploy for the build it ships, so a format bump can no longer leave a page silently declining its images. Try it.

2026-09-10examples

Parrot - talk for a few seconds, and the browser reads the poem in your voice.

▾

A third dasLLAMA example on the examples page. Press record and talk; Silero VAD ends the take when you go quiet, Kyutai Pocket TTS clones the voice from it, and the text in the box - Frost's "Stopping by Woods on a Snowy Evening" to begin with, or whatever you type - is read in that voice. The recording stays in the tab. Behind it the English Pocket TTS file shrank from 152 MB to 75: the backbone and the codec transformers as Q4_K, the flow head Q8_0, the codec encoder still inside; on our 200-sentence rig it reads 3.86 WER / 4.295 UTMOS against the q8 file's 3.91 / 4.328, and we could not hear the difference. Storywish now reads through the same form with one voice and no encoder, 65 MB. Try it.

2026-09-07site

dasllama-server is a download now - one archive per platform.

▾

The OpenAI-compatible server ships as a standalone bundle for macOS (Apple silicon), Windows x64, Linux x86_64 and Linux arm64, from the rolling dasllama-server release on GitHub. Unpack it, start the watchdog (on a Mac, the app), and open the control page at 127.0.0.1:8080. The daslang SDK's bin/daslang -jit runs the same server and tunes every kernel on your own box for optimal performance.

2026-09-05examples

dasLLAMA runs in the browser - the storyteller, compiled to wasm64, opens the new examples page.

▾

The engine behind the ladder now also ships as a WebAssembly build. daspkg release wasm compiles the storyteller - llama2.c's stories15M writing a children's tale while KittenTTS nano reads it aloud - into one 26 MB wasm64 module, and the two models arrive as prepared .dlim images, 76 MB with the two English front-end packs, the same format the native engine maps. The speech thread and the parallel kernels run on Web Workers in the page; nothing is interpreted. Chrome or Edge 133+, Firefox 134+ (memory64). Try it.

2026-09-03engine

dasLLAMA speaks - KittenTTS nano and mini, Kokoro-82M, and a text front end that is nothing but data.

▾

Three text-to-speech models serve from the same engine and the same tuned kernels as the language models, on the CPU, through /v1/audio/speech on dasllama-server and a speech studio on its control page. On an M1 Max the served 8-bit lane reads a real-time factor of 0.026 for kitten-nano (onnxruntime on the same input, 0.035) and 0.072 for kokoro-82m (torch CPU, 0.097), measured 2026-09-02. Why a game wants this, how the front end became two data packs, and the afternoon an "s" turned out to be a comma: the story.

2026-09-03engine

Speculative decoding lands on Metal - Qwen 3.8's NextN head and Gemma 4's assistant drafter.

▾

Both drafter families that ship with their models now run on any Apple GPU dasLLAMA runs on, and the depth is a per-box knob the tuner mints. Measured against a stock llama.cpp on identical rendered prompts: gemma-4-26B-A4B on an M5 Max goes 122.7 to 145.9 tok/s at depth 1 (llama.cpp 97.7 to 133.7), Qwen3.8-27B on an M4 Pro 12.7 to 14.0 (llama.cpp 11.5 to 13.8). How we measured, what we tried, and why the gain is as small as it is: the story.

2026-08-30hardware

Apple M5 support lands - text, vision and speech, tuned on the new silicon.

▾

dasLLAMA now treats the Apple M5 as a first-class target: the Metal 4 tensor units where the per-box race crowns them, the vision and speech encoders on the GPU, and the CPU lane on the M5 Max's all-compute core layout. The full M5 Max board is re-minted and on the board - gemma-4-E2B Q8_0 at 9074.3 / 161.3 tok/s (llama.cpp 7413.4 / 137.0), Qwen3.8-27B at 913.9 / 27.8 (752.1 / 25.4), pp512 / tg128 on Metal. Bringing a chip up one model at a time, and why it was fun: the story.

2026-08-18models

Qwen 3.8 27B joins the board - the new dense hybrid runs on CPU and Metal.

▾

Qwen 3.8 27B is out, and it needed no new architecture support in dasLLAMA: the 3.8 generation keeps the qwen3.5 architecture - Gated-DeltaNet recurrent layers interleaved 3:1 with gated full attention - so it loads on the existing family path as a pure scale-up. The Q4_K_M file's attention and FFN planes serve natively on the kq rails, CPU and Metal alike (the DeltaNet projections transcode to q8 at load), and the model joins the official board catalog.

On an M1 Max (pp512 / tg128, llama-bench protocol): Metal - das 147.0 / 14.5 tok/s, llama.cpp 137.6 / 11.5; CPU - das 33.2 / 5.8, llama.cpp 29.8 / 5.6. On an M4 Pro: Metal - das 127.0 / 12.8, llama.cpp 126.2 / 11.5; CPU - das 45.8 / 11.6, llama.cpp 40.4 / 11.0. Full rows for both boxes are on the board.

2026-08-17engine

Metal support lands in speech-to-text - whisper large-v3-turbo at 49x realtime on an M1 Max.

▾

Speech-to-text now serves on Metal: whisper's cross-attention and decoder run on the GPU by default on Apple silicon, and the full-GPU mode powers the new gpu category on the board. On an M1 Max, whisper large-v3-turbo transcribes a 3:19 clip in 4.0 s - 49x realtime, against 36x for whisper.cpp's Metal path and 12x for our CPU lane. The board now splits ASR into the same cpu / cpu + accel / gpu categories the LLM board has.

2026-08-10site

dasllama.io is live - news, the full ladder, and the sidecar exchange.

▾

Everything measured now lives here, community submissions included. daslang.io/dasllama.html stays the promotional scoreboard and carries official measurements only.

Latest measurements.

loading the ladder…

amber = dasLLAMA, teal = the llama.cpp run it was measured against — the same pair bars daslang.io uses. Rows compare only within one engine version.