<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
<title>dasllama.io news</title>
<author><name>dasllama.io</name></author>
<link href="https://dasllama.io/"/>
<link href="https://dasllama.io/feed.xml" rel="self"/>
<updated>2026-08-18T00:00:00Z</updated>
<id>https://dasllama.io/</id>
<entry>
<title>Qwen 3.8 27B joins the board - the new dense hybrid runs on CPU and Metal.</title>
<link href="https://dasllama.io/#n-2026-08-18-qwen38-27b"/>
<id>https://dasllama.io/#n-2026-08-18-qwen38-27b</id>
<updated>2026-08-18T00:00:00Z</updated>
<content type="html">&lt;p&gt;Qwen 3.8 27B is out, and it needed no new architecture support in dasLLAMA: the 3.8
generation keeps the qwen3.5 architecture - Gated-DeltaNet recurrent layers interleaved
3:1 with gated full attention - so it loads on the existing family path as a pure
scale-up. The Q4_K_M file&#x27;s attention and FFN planes serve natively on the kq rails, CPU
and Metal alike (the DeltaNet projections transcode to q8 at load), and the model joins
the official board catalog.&lt;/p&gt;
&lt;p&gt;On an M1 Max (pp512 / tg128, llama-bench protocol): Metal - das 147.0 / 14.5 tok/s,
llama.cpp 137.6 / 11.5; CPU - das 33.2 / 5.8, llama.cpp 29.8 / 5.6. On an M4 Pro:
Metal - das 127.0 / 12.8, llama.cpp 126.2 / 11.5; CPU - das 45.8 / 11.6, llama.cpp
40.4 / 11.0. Full rows for both boxes are on the
&lt;a href=&quot;https://daslang.io/dasllama.html&quot;&gt;board&lt;/a&gt;.&lt;/p&gt;</content>
</entry>
<entry>
<title>Metal support lands in speech-to-text - whisper large-v3-turbo at 49x realtime on an M1 Max.</title>
<link href="https://dasllama.io/#n-2026-08-17-metal-asr"/>
<id>https://dasllama.io/#n-2026-08-17-metal-asr</id>
<updated>2026-08-17T00:00:00Z</updated>
<content type="html">&lt;p&gt;Speech-to-text now serves on Metal: whisper&#x27;s cross-attention and decoder run on the GPU
by default on Apple silicon, and the full-GPU mode powers the new gpu category on the
board. On an M1 Max, whisper large-v3-turbo transcribes a 3:19 clip in 4.0 s - 49x
realtime, against 36x for whisper.cpp&#x27;s Metal path and 12x for our CPU lane. The
&lt;a href=&quot;https://daslang.io/dasllama.html&quot;&gt;board&lt;/a&gt; now splits ASR into the same cpu / cpu + accel /
gpu categories the LLM board has.&lt;/p&gt;</content>
</entry>
<entry>
<title>dasllama.io is live - news, the full ladder, and the sidecar exchange.</title>
<link href="https://dasllama.io/#n-2026-08-10-dasllama-io-is-live"/>
<id>https://dasllama.io/#n-2026-08-10-dasllama-io-is-live</id>
<updated>2026-08-10T00:00:00Z</updated>
<content type="html">&lt;p&gt;Everything measured now lives here, community submissions included.
&lt;a href=&quot;https://daslang.io/dasllama.html&quot;&gt;daslang.io/dasllama.html&lt;/a&gt; stays the promotional
scoreboard and carries official measurements only.&lt;/p&gt;</content>
</entry>
</feed>
