Word count: 18894. Estimated reading time: 89 minutes.
- Summary:
- A diary entry is provided. The release and performance of Deepseek v4 Flash is discussed. Comparisons are made with other LLMs through coding and summarisation tests. Predictions regarding AI evolution, future Apple hardware, and the impact of humanoid robots on employment are detailed. The diffusion of technology within Europe is also analysed.
Friday 7 August 2026: 00:01.
- Summary:
- A diary entry is provided. The release and performance of Deepseek v4 Flash is discussed. Comparisons are made with other LLMs through coding and summarisation tests. Predictions regarding AI evolution, future Apple hardware, and the impact of humanoid robots on employment are detailed. The diffusion of technology within Europe is also analysed.
Before we get into that though, yesterday my children finished painting the west wall white, thus concluding successfully the painting of around one hundred square metres of exterior wall. I think they did really great given their ages:
To complete a job like this over multiple days, it requires a focus and self control and willingness to see things through to when they are complete which I find lacking in most eighteen year olds, never mind much younger again. Well done Clara, Henry and Julia!
Cheap open weights AI leaps forward yet again!
You may remember that I was initially keen on Qwen3 Coder Next, it was rather slow on my ancient hardware but it did work. However I found myself thereafter mostly using Step 3.5 Flash rented from OpenRouter as it was surprisingly good at coding and agentic work, and it looks like I was early compared to most to realise this – however, then Step 3.7 Flash dropped, and it was better in every way however also twice as expensive for new input BUT now they had prompt caching implemented. Step 3.7 also emitted far fewer thinking tokens than Step 3.5, so all in all the actual cost paid dropped by about half, and I’ve found myself using Step 3.7 Flash for pretty much everything since its release as it had the best ‘bang for the buck’ from my testing i.e. Pareto optimum, and to be specific:
- It is not the most capable model by any means.
- It makes many mistakes in the code it writes.
- It can take multiple attempts to perform an edit or call a tool successfully.
- BUT if you apply multiple rounds of it checking its work it does catch 98% of the bugs and bad logic it writes and fixes them correctly.
- It is sufficiently cheap that I’ve spent a total of US$14 ever on it, and that’s despite it horsing through 160 million tokens …
- From my testing on my own actual use cases, it was the optimal cost-benefit choice of LLM for all tasks where the data it processed is public (I use a local small Gemma 4 LLM for anything processing data which isn’t already on the public internet)
- Things I really like about Step 3.7 Flash: it follows instructions well, it’s very hard to jail break it out of its system prompt, if you order it to be biased or non biased in its system prompt it does as it is told, and a 196b model is feasibly likely to be runnable on consumer hardware arriving soon, so it’s worth investing into mastering this class of LLMs as your daily driver.
Amazingly, it was only five months ago that Qwen3 Coder Next (Q3CN) landed; and just three months since Step 3.7 Flash landed. Now we have the final release of Deepseek v4 Flash, and here are those LLMs compared so you can see why everybody including me is so excited by this particular LLM release and why social media (or at least my view of it) has been jammed with Deepseek v4 Flash 0731 posts for the past week:
| Qwen3 Coder Next | Step 3.5 Flash | Step 3.7 Flash | Deepseek v4 Flash 0731 | Claude Fable 5 | ||
|---|---|---|---|---|---|---|
| Released: | Feb 2026 | Feb 2026 | May 2026 | August 2026 | June 2026 | |
| MoE weights (total-active): | 80b-a3b | 196b-a11b | 196b-a11b | 284b-a13b | 6t-a400b | |
| Max input context: | 262k | 262k | 262k | 1M | 1M | |
| Typical Openrouter input cost after prompt caching: | $0.103/M | $0.100/M | $0.053/M | $0.030/M | $3.36/M | |
| Artificial Analysis Intelligence Index: | 21.1 | 26 | 30.3 | 49.9 | 59.9 | |
| Artificial Analysis Analysis Index: | 36.2 | ? | 39.6 | 69.1 | 76.5 | |
| Artificial Analysis Agentic Index: | 8.8 | ? | 21.5 | 45.7 | 52.8 | |
| AA Omniscience Accuracy: | 15.8% | 23.9% | 25.4% | 37.2% | 61.4% | |
| AA Omniscience Non-Hallucination Rate: | 9.1% | 14.8% | 15.6% | 15.6% | 45.1% | |
| SciCode: | 32.3% | 40.4% | 40.0% | 49.9% | 60.2% |
For comparison, I placed in the final column the current best performing LLM anywhere which is Claude Fable 5. It is 112x times more expensive than Deepseek v4 Flash 0731! Until now the Pareto optimum Step 3.7 performed about half as well as state of the art – now you have something 80% as capable and for nearly half the cost of the previous Pareto optimum.
Deepseek v4 Flash 0731 is a model which at 284 billion parameters is still within the realm of near-future consumer hardware: by 2028, as you’ll see later on in this diary entry, your standard new Apple Macbook Pro from 2028 onwards is expected will include similar compute and memory bandwidth to a 2017-era nVidia Volta AI accelerator board. That should run a model like Deepseek v4 Flash well at around one hundred tokens generated per second, and maybe four thousand tokens parsed per second. That’s a good bit faster than my rented edition has been, so that’s more than fast enough for serious usage.
Between now and then, and especially as the price per token has just halved again with Deepseek v4 Flash 0731 while the capabilities took another leap forwards, it makes the most sense to rent. After this AI investment bubble bursts, I fully expect prices for renting LLMs to crash spectacularly, almost to the point of free … which may make buying local LLM capable hardware a tough ask especially if somebody invents an end-to-end cryptographically secure LLM execution engine, which I’m sure is just a matter of time. However, if your new Apple Macbook Pro just comes bundled in for no extra cost the capability to run local LLMs in the hundred of billion parameter class, then that’s a different matter again.
So, this new LLM aces the synthetic benchmarks. How does it perform in real life?
Test 1: Summarising a 19th century will
Replicating this test from last January, I had it summarise the raw transcribed will of my ancestor Walter Irvine which is early 19th century legal English full of transcription errors. Even the ancient llama 3.1 8b doesn’t do terribly at this, but at the time Claude Sonnet 4.5 did a great job. How do Deepseek v4 Flash 0731 and Step 3.7 Flash perform?
| Cloud full fat Claude Sonnet 4.5 400b? 8 bit? | Step 3.7 Flash 198b 8 bit | Deepseek v4 Flash 0731 284b 4 bit | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| (from the previous post, it was said at the time it was nearly instant) | Tokens per second ~220.4 tokens/s Token count 5987 tokens Cost $0.0091834875 Duration 27.2s | Tokens per second ~178.6 tokens/s Token count 4347 tokens Cost $0.00182941902 Duration 24.3s | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
This is a comprehensive probate document from Edinburgh dated November 17, 1825, concerning the estate of Walter Irvine, Esquire, who died on January 7, 1824, at Luddington House, Egham, Surrey. He held properties in the Island of Tobago (West Indies), Luddington House in Surrey, and Denino in Fife, Scotland. Scottish Estate Inventory (Total: £5,258 6s 2d)I. Rents and Arrears Due (£4,569 13s 2½d):
II. Bond from John M. Harlaw of Kirkton: £245 5s 1d plus interest to January 1824 (Total: £262 15s 1½d) III. Bond by Road Trustees (Anstruther-Crail to St. Andrews): £150 principal plus interest (Total: £162 7s 4½d) IV. Bond by Lundenfoot Road Trustees: £50 with compounded interest from 1821-1824 (Total: £55 1s 2d) V. Proceeds from sale of wood: £33 10s VI. Balance due from Walter Dickson WS (Writer to the Signet): £174 16s 11d plus interest (Total: £174 18s 11d) Will Provisions (dated October 18, 1823)Primary Beneficiaries:Wife - Catherine Irvine (née Gordon):
Three Legitimate Daughters (each receiving £35,000):
Natural Daughter - Betsey (of London):
Other Significant Bequests:Ann Stewart (sister-in-law and widow of Stewart):
Brother's Estate:
Nephews:
Nieces (Clara and Marian Irvine - daughters of Christopher William):
Jane Armstrong (formerly Jane Currie, wife of Francis Armstrong):
Mary (reputed daughter of a negro woman named Linda in Tobago):
Other Named Beneficiaries:
Continuing Annuities:
Trust Structure and Powers:Executors and Trustees:
Powers Granted to Trustees:
Specific Provisions for Children's Trusts:
West Indies Property:John Leith Dispute:
Plantation Management:
Scottish Property Requirements:Election Doctrine:
Residuary Estate:After all debts, legacies, and annuities paid, the residue divided equally among three legitimate daughters as tenants in common:
Elizabeth's share for her separate use with power to dispose by deed or will despite coverture. Probate Details:Proved: January 28, 1824, in Prerogative Court of Canterbury Initial Grant of Administration to:
Reserved Power for: John Hamilton, Christopher Irvine, and Thomas Bruce to apply later Witnesses to Will:
Codicil: Dated same day (October 18, 1823), adding £100 bequest to Susannah Ann Hamilton Scottish Confirmation: Presented November 17, 1825, before George Carshin Jr., Commissioner of Edinburgh, with deposition by W.R.K. Douglas confirming the inventory and will details. The document demonstrates extensive wealth with significant properties in Scotland, England, and the West Indies, careful provision for family members including an illegitimate daughter, and sophisticated trust arrangements typical of wealthy estates in the early 19th century. |
Document Date: 17 November 1825 (Edinburgh) 1. Inventory of Personal Estate (Scotland)Presented by executor William Robert Keith Douglas (WS) and signed by him and George Carfrae Jr. (Commissioner). The total value of the Scottish estate is £5,258 / 6 / 2 (£5,258.31).
The testator also owned property in England; letters of administration for that part were already taken out by the executors. 2. Deposition (25 October 1825)Executor W. R. K. Douglas swore that:
3. Probate
4. Last Will and Testament – Detailed Summary4.1 General Directions
4.2 Provision for Wife – Catherine Irvine
4.3 Remission of Debt
4.4 Bequests to Individuals
4.5 Trust for Natural Daughter Betsy
4.6 Main Trusts for Daughters Christina Charles and Catherine Grace
4.7 Investment and Management Powers of Trustees
4.8 Devise of Real Estate
4.9 Executors and Guardians
4.10 Election Regarding Scottish Estates
4.11 Covenant for Daughter Elizabeth’s Marriage Settlement
4.12 Residuary Estate
4.13 Receipts and Discharges
4.14 Powers of Executors
4.15 Appointment of New Trustees
4.16 Indemnity of Trustees
4.17 Management Powers
4.18 Execution
5. Codicil (18 October 1823)
6. Affidavits and Probate DetailsHandwriting Affidavits (for the codicil)
Attesting Witness Affidavit
Grant of Probate
7. Key Figures at a Glance
This document is a comprehensive and intricate testament disposing of a substantial trans‑Atlantic estate, providing generously for the testator’s wife, daughters (both legitimate and natural), other relatives, and servants, while establishing elaborate trusts to preserve wealth across generations. |
This document is a collection of probate records concerning the estate of Walter Irvine, an Esquire who held property in the Island of Tobago (West Indies), Luddington House in Surrey, and Denino in Fife, Scotland. He died at Luddington House on 7 January 1824. The documents date from between late 1823 and late 1825 and comprise an inventory of his Scottish personal estate, a sworn deposition, his last will and testament (dated 18 October 1823) with a codicil, supporting affidavits, and the formal grant of probate from the Prerogative Court of Canterbury. I. Inventory of the Personal Estate (Scotland)The inventory was presented in Edinburgh on 17 November 1825 by Mr. Walter Dickson, Writer to the Signet, and made out by William Robert Keith Douglas, Esquire, one of the Lords of the Admiralty and an executor under the will (which was proved in the Prerogative Court of Canterbury in 1824, under the Act 48 Geo III cap 140). I. Rents and Arrears of Rent due to the deceased
II. Balance of debt due by John Hay Harlaw, Esquire of Kirkton
III. Sum in bond by the Trustees of the Anstruther & Crail to St. Andrews road
IV. Sum lent to the Trustees of the Lauderfoot road
V. Proceeds of roupes of wood due at death: £33. 10. 0VI. Balance due by Walter Dickson WS on account current (31 Dec 1823)
Total value of the deceased's estate in Scotland: £5258. 6. 2
II. Deposition (Edinburgh, 25 October 1825)Before George Carfrae Junior, Commissioner appointed by the Commissaries of Edinburgh, William Robert Keith Douglas swore that:
III. The Last Will and Testament (18 October 1823)The will revokes all prior wills, codicils, and testamentary dispositions. It directs all just debts, funeral expenses, probate costs, and legacies to be paid promptly out of the personal estate, with the real estate charged for any deficiency. Provision for the Wife — Catherine Irvine
Remission of Debt
Bequest to Ann Stewart (sister-in-law)
Other Bequests
Trust for Natural Daughter "Betsy"
Bequest to Jane Armstrong (formerly Jane Currie)
Bequests to Nieces Clara and Marian Irvine
Bequest to Mary (reputed daughter of enslaved woman Linda)
Annuities to Others
Bequests to Nephews
Main Legacies for Daughters Christina and Catherine
Powers of Trustees (Maintenance and Advancement)
Investment Powers
Devise of Real Estate
Appointment of Executors and Guardians
Election Requirement for Scottish Estates
Covenant Regarding Daughter Elizabeth's Marriage Settlement
Residuary Estate Distribution
Receipts, Powers of Executors, New Trustees, Indemnity, Management
Execution and AttestationExecuted on 18 October 1823 at Luddington House, signed and sealed by Walter Irvine across 19 sheets, in the presence of witnesses:
IV. CodicilDated 18 October 1823, adds a bequest of £100 to Susanna Ann Hamilton, daughter of John Hamilton, Esquire. V. Affidavits and Grant of ProbateAffidavit (Codicil)
Affidavit (William Adams, subscribing witness)
Grant of Probate (28 January 1824)
VI. CertificateLondon, 12 November 1825 — William Robert Keith Douglas certifies that this is the probate of the will of the late Walter Irvine referred to in his 25 October 1825 deposition on the inventory. Key Points at a Glance
|
Scroll down to see more ...
The good news is that it was much harder for me to determine the strengths and weaknesses of each of the two models compared to Claude from last year: Claude Sonnet 4.5 definitely still wins by a long shot, but the gap is much smaller than it was for Qwen 3.0 30b. Something which piqued my interest is why the Chinese models are so keen on outputting blow-by-blow structure of the original document, and I wondered if it is an artefact of the Mixture of Experts (MoE) design. So I also tested Gemma 4 31b which is dense and Gemma 4 28b-a4b which is MoE, and indeed the same blow-by-blow structure appears for the latter. I guess that kinda makes sense? Incidentally, Gemma 4 31b did surprisingly poorly on this test, I had assumed it would beat Deepseek v4 Flash as the Gemma models are well known to be better at English language nuance than the Chinese models, and while yes it did very well at picking out the right essential points from the will, it didn’t pick enough of those essential points despite being told to be detailed. Maybe I needed to say ‘very detailed’? Don’t get me wrong, the quality of Gemma 4 31b’s output was good, but it was short and to the point as it were, and too much short and to the point in fact.
Re: our two models, I think Step 3.7 produces a better structured documents – it is keen on tables – and it is more terse than Deepseek v4 which gives too much irrelevant detail, plus it writes English better in my opinion: less fluff, more densely packed. Deepseek v4 on the other hand did cost one fifth the amount which Step 3.7 did, and it’s not that much worse. Still, Step 3.7 wins this test on quality of output, if you exclude Claude.
Claude from last year is much better at synthesising the document together e.g. it groups daughters together, it has realised one is illegitimate, it orders items in a reasonable priority for most human readers, and it has collapsed all the multiple sections from the original into the minimum possible set. The Chinese models, despite getting towards a similar 400 billion parameters of Claude from last year, have a way to go yet, assuming that they’ll ever get there as they have a MoE design.
Test 2: Analyse a code implementation of a specification and implementation plan
Last few weeks I have been working on atomic_wait() for C, which essentially
ports the same feature from C++ 20 into the next C standard – though we shall
be adding some additional APIs, as we don’t care much for the C++ API. Myself
and fellow committee member Jens Gustedt came up with a draft WG14
proposal paper over a number of weeks, then I iterated having Step 3.7 Flash write
a detailed implementation plan for a reference implementation using another
hand written reference implementation for a separate WG14 proposal as a template.
It did struggle a bit with writing out the plan, and I had to hand hold it a fair
bit, but we got there.
The single most important part of the plan file is probably this which describes when a proxy atomic must be used which is indexed via an internal hash table, or whether the atomic wait can be passed through to the platform specific API directly:
| Backend | 1 byte | 2 bytes | 4 bytes | 8 bytes | Hash table needed? |
|---|---|---|---|---|---|
Linux (FUTEX_WAIT/FUTEX_WAKE) |
✗ | ✗ | ✓ | ✗ | For 1-2-byte and 8-byte; futex is 32-bit only (int *uaddr, int val) |
macOS (UL_COMPARE_AND_WAIT/UL_COMPARE_AND_WAIT64) |
✗ | ✗ | ✓ | ✓ | For 1-2-byte, or sub-native-width types |
Windows (WaitOnAddress) |
✓ | ✓ | ✓ | ✓ | Never — all operand sizes bypass |
FreeBSD (UMTX_OP_WAIT/UMTX_OP_WAKE) |
✗ | ✗ | ✓ | ✓ | For 1-2-byte, or sub-native-width types; UMTX_OP_WAKE accepts a count parameter directly |
pthreads fallback (pthread_cond_wait) |
✗ | ✗ | ✗ | ✗ | Always — no kernel tracker exists |
✓ = kernel primitive available; hash table is bypassed. ✗ = no suitable kernel primitive; must use the user-space hash table.
So, the design’s essential points are:
- There are multiple implementation backends for each platform specific API.
- The public API is able to pass through to the kernel API directly for some or all atomic types depending on backend.
- For the remaining types, an internal hash table maps an atomic’s address in memory to its proxy atomic which IS compatible with the kernel API.
I asked Step 3.7 Flash to implement the reference library using the plan and proposal as guides. It replicated over the mildly changed parts of the hand written source template well enough, but really struggled with designing and writing the implementation. It came up with a poor architecture and lots of repeated and verbose code all of which was unnecessary. So I ripped out all of that, and hand wrote an implementation for Mac OS which has compile time replaceable backends for different platforms. I then asked Step 3.7 Flash to implement the backends for (i) Microsoft Windows (ii) Linux (iii) FreeBSD and (iv) generic pthreads using the Mac OS backend as a template.
Here it did better, though it messed up the calling protocol, swapped a few arguments around, got error handling wrong etc. All not great, but then Deepseek v4 Flash 0731 was announced, and I thought it might be more interesting to see how both models would analyse the current implementation. I asked both this prompt:
Perform a comprehensive review of the implementation, comparing it to the plan and proposal. Write out any deviations, problems, bugs or any other issue into docs/deepseek_v4_flash_0731_analysis.md. Be detailed, also be aware that there is likely a race condition in the implementation which is causing lost wakes. Try to identify the cause without debugging.
Step 3.7 Flash took about four minutes to complete its analysis, taking
fifteen steps and costing US$0.075 dollars to process 101k tokens and
output 23k tokens. Deepseek v4 Flash unfortunately took my ‘try’ literally,
and launched gdb as part of its analysis. So I refined the prompt just
for it:
Perform a comprehensive review of the implementation, comparing it to the plan and proposal. Write out any deviations, problems, bugs or any other issue into docs/deepseek_v4_flash_0731_analysis.md. Be detailed, also be aware that there is likely a race condition in the implementation which is causing lost wakes. Identify the cause without debugging i.e. do NOT run gdb, do NOT run test programs - inspect and analyse the code ONLY.
… which worked, and it took ten minutes to complete its analysis, taking sixteen steeps and costing US$0.029 dollars to process 81k tokens and output 39k tokens. That’s about 2.6x cheaper, incidentally.
Its analysis was deeper, but it missed some easy issues Step 3.7 caught:
| Step 3.7 Flash 198b 8 bit | Deepseek v4 Flash 0731 284b 4 bit | |||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Step 3.7 Implementation Analysis1. Deviations from Plan/Proposal1.1 Windows 1-byte and 2-byte operations use hash table fallbackPlan: "Windows: always bypass (all sizes 1–8 bytes). The Windows backend never uses the hash table." Actual: Files: 1.2 macOS
|
| Issue | Severity | Location |
|---|---|---|
Hash table proxy stuck at atomic=1 after notify |
Critical | atomic_wait_common.ipp.ipp:220-222 |
| Lost-wake race: notify before park in hash table path | Critical | atomic_wait_common.ipp.ipp:41-45 |
atomic_wait_generic returns -1 instead of looping on proxy error |
High | atomic_wait_common.ipp.ipp:411-418 |
| Lost-wake race: notify without store in native path | High | atomic_wait_linux.c.ipp:43-44 |
Windows notify_all uses WakeByAddressSingle |
Medium | atomic_wait_windows.c.ipp:100 |
| Windows 1/2-byte ops use hash table instead of bypass | Medium | atomic_wait_common.ipp.ipp:524-528 |
wait_on_address32 discards errno |
Medium | atomic_wait_linux.c.ipp:50-51 |
macOS notify_all ignores max_threads_to_wake limit |
Medium | atomic_wait_macos.c.ipp:99-106 |
| No tests for hash table path | High | test/atomic_wait_test.c |
| No tests for notify-without-store | High | test/atomic_notify_test.c |
The most severe issue is the hash table proxy state machine (section 2.1 / race 3.1). The proxy is a single-slot binary semaphore that is set to "notified" (1) by the notifier but never reset to "waiting" (0) after consumption. This makes the hash table path fundamentally broken for any object that receives more than one notify, or for any waiter that arrives after a notify has already been issued for that object.
Deepseek v4 Flash (0731) — Implementation Review
Date: 2026-08-01
Scope: Reviewed the wg14_atomic_waits implementation against
docs/proposal.md (the authoritative spec) and plans/wg14_atomic_waits-reference-library.md.
Method: Static inspection of every include/ backend, src/ wrapper, the public
header, config.h, the tests, the build files, and CI. No debugger was run and no
test program was executed; all conclusions come from reading the code.
1. Executive summary
The implementation is structurally faithful to the proposal (all six public APIs exist, width dispatch for 1/2/4/8-byte atomics, native wait/notify backends with a hash-table fallback), but it contains one primary correctness race that produces lost wakes (Section 2), a broken pthreads backend (Section 3), a timeout smuggled in as an error on Linux (Section 4), a Windows lost-wake for multi-thread wakeups (Section 5), plus a number of smaller deviations from the plan/proposal.
The single most important finding: the per-object "proxy" used by the hash-table path sets a 0/1 notification flag that is never cleared while the wait-queue node is alive. This is the lost-wake/livelock race the task asked to identify.
2. PRIMARY RACE: the proxy notification flag is never reset → lost wakes / busy-spin
Location
include/wg14_atomic_waits/detail/impl/atomic_wait_common.ipp.ipp:
..._WAITmacro (lines 37–38):wait_on_address32(&(x)->atomic, 0, (timeout)) /* wait while proxy->atomic == 0 */..._WAKEmacro (lines 39–45):atomic_store_explicit(&(x)->atomic, 1, release), /* mono-directional: 0 -> 1 only */ wake_by_address32(&(x)->atomic, max_threads_to_wake)atomic_wait_generic()(lines 327–435) — the shared "park by proxy" loop.atomic_notify_generic()(lines 437–456) — the shared "set flag + wake" path.
The bug
A waiter parks by calling WAIT(item,...) which blocks while item->atomic == 0
(i.e. FUTEX_WAIT(&item->atomic, 0)). A notifier sets item->atomic = 1 and wakes.
There is no code anywhere that ever writes item->atomic back to 0 while the
wait-queue node is alive. The only place it is reset is at node creation inside
hash_table_find_or_create() (lines 220–222), which happens only when a brand-new
proxy_waiter_t is allocated. A node is freed only when use_count drops to zero.
Consequence — two interleaving outcomes
(a) Re-park after a wake never sleeps (livelock). Once any notify has fired on a
node, item->atomic is stuck at 1 for as long as the node lives. Any waiter that is
woken and must re-park — the proposal explicitly requires re-park on spurious wake, and
the code implements it as the top of the loop — calls
FUTEX_WAIT(&item->atomic, 0) while the value is already 1. The kernel compares
1 != 0 and returns EAGAIN immediately. Every subsequent iteration of the loop
returns immediately, so the waiter never sleeps again; it degenerates into a tight
100%-CPU busy-spin for the whole remaining lifetime of that node.
(b) Wakeups are lost because there is no sleeping thread. Because (a) means waiters
stop sleeping, a later genuine producer store + notify_* sets the flag (already 1)
and issues FUTEX_WAKE, which has nothing asleep to wake. The notify is therefore
effectively lost for the purpose of the sleep/wake contract; correctness then depends
entirely on the busy-spin poll observing the value change, which is not the semantics the
proposal defines and not what a correct reference implementation should do.
Why the analogous futex idiom would be safe but this one is not
The correct pattern guards the "am I allowed to sleep" decision on the same state that the notifier flips, and the notifier re-arms the state before waiting:
- waiter:
s = counter.load(); if (value == expected) futex_wait(&counter, s); - notifier:
counter++; futex_wake(...)Herecounteris a strictly increasing generation so the waiter can always detect a change that happened between its load and its sleep. The implementation instead uses a single 0/1 flag that is never re-armed, so the invariant "atomic == 0⇔ a notify is pending/expected" is destroyed after the first notify.
Which configurations suffer
This path is the fallback for every backend whenever the operand cannot be handled directly by the kernel primitive, i.e. exactly the cases the proposal/plan force through the hash table:
- Linux: 1-, 2- and 8-byte operands (
HAVE_WAIT_ON_ADDRESS_*is 32-bit only). - macOS / FreeBSD: sub-native widths (1/2-byte).
- pthreads backend: every operand (there is no kernel per-address waiter).
The 4-byte Linux/macOS/FreeBSD/Windows fast paths and the atomic_wait_expected
native-width path bypass the proxy and are not affected by this flag, but the
8-byte-on-Linux case — a perfectly legal and likely test target — is affected.
Recommended fix direction
Replace the 0/1 flag with a monotonically increasing sequence number that the waiter reads before parking and passes as the futex compare value, and that the notifier increments before waking. Reset-on-rearm must happen on the waiter side before the sleep decision, under the same lock used to re-check the object value (or rely on the kernel re-check for the object value itself as the futex fast path already does).
3. pthreads backend is fundamentally broken (hangs / lost wake)
include/wg14_atomic_waits/detail/impl/atomic_wait_pthreads.c.ipp:
..._WAIT(lines 45–46) ispthread_cond_wait(&(x)->atomic, pthreads_mutex()).pthreads_mutex()(lines 61–73) returns a_Thread_localmutex, i.e. a different mutex object per thread.
Problems:
pthread_cond_waitrequires the passed mutex to be held by the calling thread. Inatomic_wait_genericthe waiter has released the hash-table lock (line 387) and then enterspthread_cond_waitwith a mutex that is never locked. This is undefined behavior; on glibc it typically fails immediately (EPERM) so the wait “succeeds” without ever blocking — again a busy-loop — and there is no guarantee the node is protected.- The broadcast hand-off is not protected by the mutex the waiter sleeps on. A
notifier holds the global hash-table lock and calls
pthread_cond_signal(via the..._WAKEmacro, lines 47–53). The classic lost wake occurs when the notifier signals between the waiter’s re-check (value still equal toexpected, line 380) and itspthread_cond_wait: the signal is dropped and the waiter blocks forever. With a futex, the kernel’s value re-check/EAGAIN saves this; withpthread_cond_tthere is no such guard and there is no predicate/flag protecting the check, so the wait is a genuine, permanent lost wake (a hang). - Even the
INIT/DESTROYmacros treatpthread_cond_tthrough the genericproxy_waiter_t.atomicslot, but the sharedatomic_wait_genericstill performs flag-style logic (settinguse_count, etc.) that is meaningless for a condvar.
Because CI runs ALWAYS_USE_PTHREADS_BACKEND=ON on Linux and macOS, this path is
exercised, but its crashes/hangs are exactly the class of lost-wake bug being reported.
4. atomic_wait_expected mis-reports a timeout as an error on Linux
include/wg14_atomic_waits/detail/impl/atomic_wait_common.ipp.ipp,
atomic_wait_expected_32() (lines 595–666), plus the Linux
wait_on_address32() (atomic_wait_linux.c.ipp lines 37–52).
- The Linux
wait_on_address32returns0on success/EAGAIN/EINTR and-1on any other error, not-errno. A genuine time-out ofFUTEX_WAITtherefore comes back as-1(witherrno == ETIMEDOUT). - The caller’s error branch (lines 651–659):
if(ret2 < 0) { if(duration != NULL && ret2 != ETIME && ret2 != ETIMEDOUT) { errno = -ret2; return -1; } }ret2is-1, which is never equal to the positiveETIME/ETIMEDOUTconstants, so the condition is always true for anyret2 < 0when a duration was supplied. A clean time-out returns-1(error), not0(time-out) as the proposal requires:Returns: … returns zero … or duration timeout occurs.
This is timing-dependent — if the pre-wait clock_gettime check (lines 631–637) happens
to notice expiry first it returns 0 cleanly — so the 1 ns test in
atomic_wait_expected_test.c is flaky, but the underlying error path is wrong.
5. Windows wake_by_address* only ever wakes a single thread → lost wake
include/wg14_atomic_waits/detail/impl/atomic_wait_windows.c.ipp,
wake_by_address32/wake_by_address64 (lines 92–124):
if(WakeByAddressSingle((PVOID)(uintptr_t) object)) return (max_threads_to_wake == 1) ? 1 : 1;
return 0;
- The
max_threads_to_wakeparameter is ignored; bothatomic_notify_allandatomic_notify(..., max_threads_to_wake=N>1, ...)call this and wake exactly one thread viaWakeByAddressSingle. The correct routine formax != 1isWakeByAddressAll. Every other waiting thread is left parked → lost wake. - The
(max_threads_to_wake == 1) ? 1 : 1ternary is dead code (both branches are1).
This makes the Windows backend incorrect for atomic_notify_all and for
atomic_notify with max_threads_to_wake > 1, which the plan marks as mandatory
behaviour.
6. macOS timeout conversion deviates from the plan
atomic_wait_macos.c.ipp, wait_on_address32/64 (lines 55–68, 77–90):
- The plan (Step 11) requires:
*duration→ nanoseconds, cap eachulock_waitcall atUINT32_MAX(~4.29 s) and loop for longer durations. - The implementation instead converts once to microseconds
(
tv_sec * 1000000U + tv_nsec / 1000U) and passes it in a single call with no cap and no loop. For any duration ≥ ~4295 s the microsecond value overflowsuint32_t, and durations beyond ~4.29 s are not split across multiple calls, so the accumulated wait can be far shorter than*duration— violating the proposal’s “total accumulated time … shall be at least*duration”. - The code also declares private
extern __ulock_wait/__ulock_wakeinstead of including<bsd/sys/ulock.h>as the plan directs; functional risk if SDK/version behaviour differs.
7. FreeBSD 8-byte UMTX_OP_WAIT argument-order inconsistency
atomic_wait_freebsd.c.ipp:
- 4-byte:
_umtx_op(object, UMTX_OP_WAIT_UINT, expected, (long)&umtx_time)(lines 57–58) — passes the expected value in the value slot. - 8-byte:
_umtx_op(object, UMTX_OP_WAIT, (long)&umtx_time, (long)expected)(lines 88–89) — passes the timeout pointer in the value slot and the expected value in the address slot, i.e. the two are swapped relative to the 4-byte call.
The two calls are internally inconsistent, so at least one passes the operands in the
wrong order; UMTX_OP_WAIT (8-byte) is almost certainly wrong and will compare against
garbage / misbehave.
8. Return-value deviations from the plan
atomic_notify_32(lines 668–691) returns1 + reton a successful CAS, whereretis the number actually woken (0 on the proxy path when no node exists, 0 on the futex path when nothing is parked). So CAS-success-with-no-waiters returns1(positive), whereas the plan explicitly states:atomic_notify… Returns 0 if the CAS fails or no waiters are parked.(The proposal’s “possibly one plus the number woken” makes
1defensible, so this is a plan deviation, not a proposal violation — noted for completeness.)atomic_wait_genericsetsret = 1(line 388) before the first actual park call. The proposal ties “positive” to “suspended at least once”; the intent-to-park flag is acceptable but slightly loose.
9. Width-dispatch macros silently no-op on unsupported widths
atomic_wait.h, _WG14_ATOMIC_WAITS_IMPL_atomic_wait* / _notify* (lines 107–201):
- Each macro is
if (sizeof==1) … else if (==2) … else if (==4) … else if (==8) …with no finalelse(anddo{}while(0)). A_Atomictype of any other width (e.g. 16 bytes, or along double, or a 0-width type) compiles to a silent no-op rather than a compile-time error, which can mask misuse. The plan states widths 1/2/4/8 are supported; a diagnostic would be safer.
10. Header-only / ODR notes
- The
atomic_wait_*_N/atomic_notify*_Ndefinitions in the.ippfiles are not themselves markedWG14_ATOMIC_WAITS_INLINE/static(only their prior declarations inatomic_wait.hcarryWG14_ATOMIC_WAITS_EXTERN, which isinlineonly whenWG14_ATOMIC_WAITS_ENABLE_HEADER_ONLY). In the multi-TU header-only build (header_only_test*.cpp) this relies on the C inlineextern-inlinemodel. It is workable on GCC/Clang but fragile; the plan required every backend helper to be explicitlystatic inlineto guarantee ODR safety. hash_table()usesWG14_ATOMIC_WAITS_IGNORE_MULTIPLE_DEFINITIONS(weak), so the singleton table is fine across TUs; this part is OK.
11. Smaller issues
errnonot preserved on success/failure paths ofatomic_wait_expected_32: it writeserrno = ETIMEDOUT(line 636) anderrno = -ret2(line 656), clobbering the caller’serrnoeven thougherrnopreservation was a stated plan requirement (Steps 8–12). The futexwait_on_address32does restoreerrno, but the caller overwrites it afterwards.EINTRtreated as success (return 0) inwait_on_address32(Linux line 45) rather than retrying; a signal during a timed wait silently shortens it.atomic_notify_genericdrops a notify when no node is registered (find_or_create withincrement_use_count=false, lines 444–445). For the futex backends this is masked by the waiter’s under-lock re-check; for the pthreads backend it is not masked (see Section 3), which is where the drop becomes a hard lost wake.- Hash-table growth/deletion re-probes correctly (cluster-shift deletion, lines 271–
280), but
hash_functruncates a 64-bit pointer to 32-bit before mixing (lines 108–113); not a correctness bug, just avoidable collisions.
12. Structure deviations from the plan (non-bug)
- The plan (Step 7) specified that
atomic_wait/atomic_wait_explicitbe thin wrappers delegating toatomic_wait_expected, andnotify_one/notify_alldelegate to awake_waitershelper. The implementation instead uses per-widthatomic_wait_1/2/4/8+ direct kernel calls for the native widths, with the hash-table path only for the fallback. This is a reasonable and arguably more correct choice given thatatomic_wait_expectedis native-width-only, but it is a deviation from the plan’s described structure. benchmark_atomic_wait_test.cdoes not actually run a benchmark (returns 0 after declaring an unused consumer), so CI excludes it correctly but it validates nothing.
13. Conclusion
The implementation gets the public API shape, the width dispatch, and the native-width
futex path right, but the hash-table/fallback proxy logic has a one-way notification
flag that is never re-armed, which is the race responsible for lost wakes: waiters that
re-park on a node that has already been notified stop sleeping and busy-spin, and
subsequent notifications have no sleeping thread to wake. On the pthreads backend the same
logic is additionally broken by a thread-local mutex and a missing predicate under
pthread_cond_wait, producing hard lost-wake hangs. These are compounded by a Linux
timeout→error bug, a Windows single-wake bug for notify_all/max>1, a macOS timeout
conversion that violates the minimum-duration guarantee, and a likely FreeBSD 8-byte
argument swap. All were identified by code inspection only.
Scroll down to see more ...
Differences:
- Only Step 3.7 noticed that the Windows backend doesn’t bypass for all atomic sizes as it should according to the implementation plan.
- Step 3.7 false positives ‘atomic_notify_generic creates a proxy even when no waiters exist’.
- Step 3.7 side mentions the lack of a generational count: Deepseek v4 gets properly into the issue and describes it in full and does it well at that.
- Deepseek v4 has a much better handle on what’s broken in the pthreads backend and why.
- Only Deepseek v4 spots the inverted arguments in the FreeBSD backend.
Also, in general, the Deepseek v4 report reads more coherently and gets into the big issues quickly and tersely, whereas the Step 3.7 report is bitty and kinda all over the place.
Neither did a good job of identifying where the implementation plan or the implementation deviate from the proposal. The implementation plan specifically states at its top:
docs/proposal.mdis the authoritative spec. Behavior, return values, and memory-order semantics must match it exactly.
After this I stopped using Step 3.7 and exclusively used Deepseek v4. Perhaps the latter was overwhelmed by all the defects and deviations from specification with this first analysis which is why it didn’t perform well – all I can say is that later on, perhaps as the implementation’s major bugs got fixed which made logic analysis easier, Deepseek v4 Flash began to seriously impress me with its analysis capabilities. One still has to go through multiple rounds of something like:
Exhaustively compare the implementation to the proposal, looking for all cases of deviation, bug, issue, concerns and corner case where the implementation does not match the proposal, or the proposal may not be implemented in full. Be very detailed, thorough and methodical in your approach - go that extra mile each and every time. Analyse in depth all implementation backends and all combinations of build configuration options, also analyse in depth all error handling and control flow paths not normally taken. Update plans/combined-analysis with your findings.
… and then you rinse and repeat iterations of that, fixing one by one all the things it finds, and doing so until it finds nothing important in your opinion. So in that sense it’s like Step 3.7, but where it massively improves is on one shot fixes for CI failures: you simply tell it which CI failed and copy and paste the failure text. It then had a 100% success rate at finding and fixing the CI failure even for platforms it could not debug locally – it did so simply by inspection and analysis, including inspecting online the kernel sources for Linux, FreeBSD or ReactOS (to get an insight into Windows).
In fact, there was an especially impressive bit where it found that
Apple Clang 17.0 only would produce invalid binaries if symbol visibility
was set to hidden and a specific tail call optimisation caused a
function to be inlined into main(). It went off decompiled the Apple
Clang binary, compared it against the LLVM clang source code, found
the exact bit of problematic reverse compiled source code, wondered to
itself it it ought to patch the Apple Clang binary, and eventually decided to
instead hack around the problem for this specific instance and it even
added an informative explanatory comment to say why its hack was there.
Now, I’d read of Claude doing stuff like that. I’d seen Step 3.7 analyse
the assembler in binaries to figure out why they weren’t performing
as expected. But to actually localise a bug in a third party precompiled binary via reverse engineering?
That was new to me. No doubt it did take rather a long time to do all
that – due to Deepseek v4’s immense popularity right now, it has not
been running quickly, as little as 30 toks/sec. But I could leave it
chug away on its own safely I found, whereas Step 3.7 had a nasty
habit of occasionally wrecking your git repo or going off and installing
huge bits of software it didn’t really need via brew.
After the WG14 atomic waits reference library was finished, I had spent US$5.02 on 465 million tokens. That is US$0.0108 per million tokens. Yes that is an awful lot of tokens – in fact, I have consumed 913 million tokens ever on OpenRouter, so this one project consumed half my lifetime total – but the quality of implementation created is very high in my opinion. I would estimate it would have taken me over one hundred hours to create a similar quality implementation by hand – instead this cost me less than ten hours in total, which was almost entirely spent reviewing its work and giving direction on what to do next. Five dollars for ninety hours of my life back to do more interesting work is a bargain.
Test 3: Subjective experience of using each LLM to get work done
I think it’s fair to say I’ve been repeatedly wowed by Deepseek v4 Flash 0731’s capabilities. I HAVE found that you should not let it take architecture direction decisions: always ask it to present a menu of implementation options, and you’ll find half the time its recommended implementation is the wrong one. So that part sucks. But when you choose on its behalf the right implementation option, 98% of the time it does a great job: it matches the style and form of the existing codebase, it avoids writing copy and paste code and instead hoists common routines into reasonable locations in reasonable common header files, and the code quality written is well above most of the programmers I’ve ever worked with, with only very occasional slip ups. I really like the much improved one-shot fix capability, especially for platforms and architectures I can’t run on my system where I have the LLM agentic harness running. That’s been a HUGE timesaver: no more having to boot up Windows VMs etc to diagnose some random failure on CI.
I very much like its performance analysis. I asked it to make this codebase go faster. It spent some minutes pondering and reading code, and it told me we ought to use triangular probing instead of quadratic probing in the open addressed hash table as the buckets are a power of two, so the triangular probing would ensure better scattering of entries avoiding collisions. It one-shotted the new implementation, then benchmarked the difference, then twiddled a few unrelated items by parsing through the optimised disassembly as it knew my main ask was for improved performance. It then spat out hard benchmarks: 470 nanoseconds reduced to 50 nanoseconds. Impressive. It also generated comprehensive tests for scalability under load and that bucket growth did work perfectly under heavy multithreaded load. Even more impressive.
I asked it to add the Fil-C toolchain to the CI (this is a guaranteed memory safe C/C++ toolchain). It went off and found the documentation on the web, followed the instructions, set up the appropriate Github CI actions, adjusted the codebase where necessary as the Fil-C libc is musl rather than glibc, then to test it it installed a Linux VM as this is a Macbook, installed the correct AArch64 edition of Fil-C rather than the x64 one the CI uses, and ran the test suite via the local Linux VM under Fil-C. Worked first time in a single shot too – it didn’t make a single mistake. Had I done that by hand, I definitely would have made a mistake at least once – I know from past experience that setting up the Fil-C toolchain is finicky.
I originally had asked Step 3.7 to create this new reference library using an existing hand written reference library as its template. It didn’t do too well at that, and with hindsight I wish I’d have wiped what it did and started from scratch with Deepseek v4. Deepseek v4 Flash makes far fewer mistakes within the test harness (Kilo Code) and doesn’t need to self correct anything like as much. It also gets tool calling in the Kilo code harness right almost all of the time, unlike Step 3.7. I suspect it would have done a better job at mimicking the template into this new library, but I guess I’ll not find out until the next time I write a new reference library for WG14.
Deepseek v4 Flash is much more prone to proactively fix bugs and issues
without being explicitly told it can first. It seems to ask for forgiveness
rather than permission. So you’ll need to be careful to always git commit
before asking it a question, otherwise it may decide your question demands
code changes. At least it doesn’t like to git reset --hard
as Step 3.7 Flash was keen on doing when it got the codebase into
a confused state, rather it properly uses git stash so any working
tree changes can be recovered.
The 1M max context of Deepseek v4 Flash makes a BIG difference! I was used to beginning to sweat as the 260k context limit approached, trying to get it to write out todo lists into Markdown files for the next session clear as I always found context compaction just didn’t work well with Step 3.7. With Deepseek, I can just relax and let it trundle on – you are still wise to start a new session from time to time as the long contexts slow its execution down, but now you can take your time about it, and more importantly, if it goes off on a long extended think or ponder or diagnosis of something you can just ignore it because it won’t suddenly run out of context. This is the first occasion I can go do some other task while it runs in the background and I don’t need to stress constantly checking its progress. Very nice!
And finally, I really like how much cheaper it is. On Openrouter comparing usage and spend from before to after, I issued twice the requests and spent half as much money. I didn’t have to babysit this model as much as before. All in all, this is my new favourite Pareto cost-benefit optimum LLM choice. Well, at least for the next three months, if the pattern so far this year continues to hold true!
Which brings me onto …
Where LLMs and AI probably are going next
We now have enough history of LLM evolution to be able to predict with reasonable reliability where things will go next. As I mentioned above, I’ve chosen a new LLM coding assistant every three months on average this year. Shall I continue to do so?
I had Deepseek v4 Flash go off and scrape the AA Intelligence Index for a spread of LLMs over the past two years off https://artificialanalysis.ai/, then plot those using a contour map:
This has contour bands for ventiles in a LLMs AA Intelligence Index score and it shows that:
- For a <= 10 billion parameter model which is feasible for me to run locally given my ancient hardware, we broke into the tens around January this year, and we would expect to break into the twenties any time around now, followed by the thirties around January 2026, and the forties around Summer 2027. So, by Summer 2027, a 10 billion parameter model will score as well as Deepseek v4 Flash does. And it’ll run well even on this old Apple M3 based laptop.
- For a <= 300 billion parameter model which is likely to run well on near future Apple Macbook Pros, we broke into the twenties around last November, into the thirties last March, into the forties last week, and we would expect to break into the fifties before the end of 2026, then into the sixties by March 2027. Reminder: Claude Fable 5 scores sixty-one. So, before the end of Spring 2027, that six trillion parameter model will score similarly to a 300 billion parameter model!
Let’s graph time directly against AA Intelligence Index score:
What strikes you about this graph is firstly by how much the open weights models are catching up with the closed weights ones – I would be surprised therefore if the Chinese government continue to release their frontier models with downloadable weights in the near future. Secondly, there is a clear structural break between 100-200b and 200-500b models – there is wide space between their trend lines. Indeed, right now 200-500b is outperforming 500-1T, which is surprising. Less surprising is the 1T+ category which has a trend line matching that of the closed weight models.
The current most intelligent LLM anywhere by this index as of August 2026 is Claude Opus 5 with a score of sixty-one, followed by the current most intelligent free to download LLM which is Kimi K3 with a score of fifty-seven (Kimi K3 is a cool 1.4 Tb of download, and generally you need about as much VRAM as the download size to run it, so that would be very expensive to run locally given current RAM prices). Deepseek v4 Flash 0731 is 167 Gb to download, so it would run well on a machine with 256 Gb of VRAM, and it gets an intelligence score of 49.9. Of course, a single score is an average, and some models are strong and weak on specific domains compared to others. https://artificialanalysis.ai/ lets you compose comparisons of LLMs, so I chose these six as representative for this discussion:
Scroll down to see more ...
As you can see, the AA Intelligence Index score is made up of lots of separate indices, each of which is then weighted into an average overall score. On some specific domains e.g. GPQA Diamond or r3-Banking, already it’s a wash between recent models. On some others, there is a clear pattern of the gap rapidly narrowing, however there are always going to be things at which a five trillion parameter model will beat the pants off a 500 billion parameter model: tool use and logic aren’t those, but specialist knowledge and reasoning will be.
In other words, yes while the AA Intelligence Index score will improve over time for smaller models, that will be in those parts of the index which aren’t specialist knowledge and reasoning. Small models simply can’t store as much knowledge as larger ones, so I expect my estimations above to be rather optimistic i.e. the score improvements will be less for the smaller parameter models than one would currently predict by extrapolation from recent past. Especially because the AA Intelligence Index score is out of one hundred, so as models max out all parts except the specialist knowledge and reasoning, they will end up running into an upper bound where only more parameters can improve some of their domain specific scores.
ALSO all this prediction is contingent on the AI investment bubble continuing to inflate. One gets cleverer small models by investing more compute into training fewer parameters. For that, you need more and cheaper compute, and for that you need to keep investing those billions. This year they think ~US$900 billion has been invested in AI, and for next year we are on track for US$1.4 trillion dollars in 2027. Even the great wealth and income of the tech multinationals will struggle to fund so much debt – even now, total free cash flow for several of them is below their debt servicing costs for the debt they’ve taken out. But that’s another diary entry. In any case, it is hard to believe that the AI investment bubble won’t pop soon, and then we’ll have whatever compute has been built out by then and it’ll only grow linearly rather than exponentially after that, much like with the late 1990s telecommunications infrastructure investment bubble.
Linear compute growth does still enable model improvements, and unsurprisingly I’d expect them to stop improving exponentially and start improving linearly instead. As with the end of Moore’s law, the slowdown will affect the biggest highest end models first, and the smaller lower end models will see a long run of continuing exponential growth before that eventually also peters out. It’s been the same with CPUs: at the very cheap end, they’ve been continuing to exponentially improve the value per dollar cost for decades after the high end went into linear improvements. I think the same will apply to LLMs: after all, if training cost per dollar goes from exponential to linear improvement, the lessons learned from making the high end a little better should translate into larger improvements at lower ends, same as for CPUs.
I find this prediction of the future FAR more believable than predictions of imminent Technological Singularity which have started doing the rounds again. I covered that in the unpublished book I wrote after St. Andrews: the Singularity is purely the result of an artefact of human perception where we tend to weigh more recent big leaps forward as more important, as they are more important to us personally but aren’t really in the bigger picture of things outside humanity. Elon Musk had an interview with the Economist week before last where he was banging on about the Singularity. I suppose that suits his purposes to market that philosophy aggressively so fewer think about seizing some of his trillion dollars of personal wealth, but I also got the impression from the interview that he actually genuinely believes that a Singularity will happen at some point. I’ll categorically state right now: no Technological Singularity will happen in my lifetime unless some very new technology turns up. Certainly nothing about Large Language Models as presently designed and implemented is capable of generalised artificial intelligence i.e. AGI. Right now we’re in the exponential growth phase because we’re pouring exponential amounts of capital in – cut the constantly increasing capital investment flows and you can say good bye to exponential LLM capability improvements, as I just described above. All that said, the near term improvements to consumer hardware WILL be significant to Economic Total Factor Productivity as the gains from this technological advancement begin to diffuse widely throughout society.
Near future hardware
So that brings me onto the near future consumer hardware to run these things. Recent leaks say that Apple have started to design their next M-series and A-series chipsets to have better than the usual trendline of improvements to compute and memory bandwidth, so your 2028 Apple Macbook Pro should locally run < 500 bn parameter LLMs quite well indeed. Here are the current rumours and leaks in a single table and graph as created from the table by Deepseek v4 (at which it was surprisingly poor at doing interestingly, I really had to poke it hard and repeatedly to generate correct looking SVG, despite it amazing performance at graph building shown above – maybe the HTML table input upset it?):
| Model / Year | Edition | GPU Cores | Memory Bandwidth | Remarks |
|---|---|---|---|---|
| Apple M3 2023 | Pro | 18 cores | 154 Gb/sec | Unfortunately my personal Macbook is the M3 Pro, the worst for running LLMs of any of the Pro Macbooks 🙁 |
| Max | 40 cores | 410 Gb/sec | ||
| Apple M4 2024 | Pro | 20 cores | 273 Gb/sec | Added a memory cache shared between CPUs and GPUs like AMD's Infinity Cache for its GPUs. This greatly improved latency. |
| Max | 40 cores | 546 Gb/sec | ||
| Apple M5 2026 | Pro | 20 cores | 307 Gb/sec | First with hardware matrix multiply and accumulate (= nVidia 'tensor cores'). LLM input parsing is approx 4x faster than M4 as a result. 1024 FP16 FMAs per core per cycle enables 70 FP16 TFLOPs for the Max edition. |
| Max | 40 cores | 614 Gb/sec | ||
| Apple M6 2027? | Pro | 32? cores | 512? Gb/sec | Expected move to LPDDR6 standard memory architecture featuring a wider 24-bit channel layout (shifting to 384-bit Pro / 768-bit Max buses) to achieve a projected 1.67x generational leap in raw memory speeds. There will be no Max nor Ultra edition of the M6, this suggests that the core will be very similar to the M5 and they only upgrade the memory bandwidth. |
| Max | 64? cores | 1024? Gb/sec | ||
| Apple M7 2028? | Pro | 48? cores | 800? Gb/sec | Rumours say the Max variant can be fitted with up to 768 Gb of RAM in your standard Macbook laptop chassis. Obviously so much RAM will be VERY expensive as Apple likes to charge steeply for additional RAM. It would be surprising if TFLOPs don't double due to implementing 2048 FP16 FMAs per core per cycle. |
| Max | 96? cores | 1600? Gb/sec |
If you extrapolate out the numbers, the Apple M7 Pro should have the same memory bandwidth as a nVidia Volta enterprise AI accelerator from 2017, and the M7 Max should have the same memory bandwidth as a nVidia Ampere AI accelerator from year 2020. Chances are that the Macbook Pro and especially Max will have more VRAM (or equivalent, see below), but in terms of compute with 1024 FP16 FMAs per core per cycle they should pretty much match a Volta and Ampere exactly: the Volta maxed out at 125 FP16 TFLOPs and the Ampere 312 TFLOPs. Both had hardware matrix multiple and accumulate, same as the Apple M-series from the M5 onwards. As mentioned in the table above, it would be surprising if the M7 doesn’t implement at least 2048 FP16 FMAs per core per cycle given that today’s nVidia Rubin chipset can do 16384 FP16 FMAs per core per cycle, and Apple tends to follow closely whatever architecture choices nVidia makes – the M5 chipset’s GPU looks awfully like a nVidia GPU, just less wide. This architectural closeness is why LLM software support tends to be nVidia first, then Apple, then AMD (which is architecturally different), then Intel (which is architecturally different again). And why LLM software support on Apple is first class, whereas although support for AMD has improved enormously, it remains second class.
There are zero rumours about this next bit, so it’s probably wrong, but I would wonder if Apple would fit so much DRAM when flash mounted as Storage Class Memory (SCM) is (i) cheaper and especially (ii) much less drain on battery life. There is zero good reason why LLMs are stored in DRAM other than there isn’t an easily available cheaper substitute, but somebody big like Apple could simply fit NAND flash where the DRAM goes. You might only write that flash with an updated LLM every few months so its endurance won’t matter, and NAND flash if mounted like RAM is nearly as fast as DRAM. As the LLM model weights aren’t mutated in RAM, this could save easily 80% of the RAM demands of a LLM, so you get to run your 200 Gb sized LLM in 40 Gb of DRAM and probably less if you shrink the size of the KV cache which is very doable if you have Ampere levels of compute on tap.
Obviously that’s pure speculation, but I do know that DRAM is hard on battery life as it must be continually refreshed. Storage class memory would be easy to fit for somebody big like Apple and it would fix the battery life impact problem. I guess we’ll find out in 2028. In any case, you would expect parsing of around four thousand tokens per second, and generation of a hundred tokens per second on the M7 Pro – and double that for the M7 Max. That’s very acceptable for an ultrabook sized laptop.
Diffusion of local LLM capable hardware throughout society
Lots of ink both physically and virtually has been spilled lamenting how Europe isn’t keeping up with the US and China on AI advancement: we aren’t investing in the electricity supply for datacentres, nor in AI research past a small fraction of what the Americans and especially the Chinese are doing. It is therefore claimed that Europe will be left behind, and left at a significant disadvantage to the US and China.
This kind of claim has been made many times before on many topics of industrial, social and political comparison between the three superpowers – and it is true that especially recently Europe has felt on the back foot as it gets bullied simultaneously by the other two superpowers, which it isn’t used to historically. However, something less appreciated is that Europe is surprisingly good at diffusing more quickly and completely the gains of an advancement than the other two superpowers: it ‘buys in’ the advancement cheap, then mass disseminates it.
That will need explaining, so to simplify: Europe, due to its unique configuration of highly competitive constituent arms length states with huge size variations, tends to diffuse innovations faster and more broadly than America or China does. This is surprising on first inspection, but think of it this way: if Ireland obtains a large current account surplus by diffusing US sourced innovations widely across its economy, all cash strapped countries elsewhere in Europe start paying rapt attention and will try to duplicate and/or improve upon whatever Ireland is doing. Ireland gets a lot of stick internationally for being a tax haven and washing the profits of US multinationals of their tax obligations elsewhere – all of which is fair – but less appreciated is that all those US multinational operated subsidiaries in Ireland do genuinely diffuse US innovations throughout the Irish economy much quicker and more completely than they could in the US where they are nowhere near as relatively economically dominant. Same goes in Switzerland and Belgium incidentally.
Obviously I’m exaggerating a touch there – at times I do wonder about diffusion of best practices in Ireland – but my point is that in superpowers such as the US and China, practice of best practices tends to be concentrated in specific economic clusters such as New York or San Francisco-San Jose in the US, or Shenzhen-Guangzhou or Shanghai in China. Whereas Europe’s economic clusters are more geographically distributed and numerous in a unique three spoke configuration:
These are the famous blue, golden and green ‘bananas’ of European economic cluster (source). Unusually they all connect together through the North of Italy, which is exactly why while the Covid pandemic may have originated in China, it turned into a global pandemic in the North of Italy as that is the most connected place to other places in world bar none other, so all global pandemics will always spread worldwide from there. As with infectious diseases, so does the global diffusion and spread of new ideas and best practices all originate from Northern Italy.
And the same will undoubtedly apply to the mass adoption and use of LLMs: the US may design the hardware and the Chinese may manufacture the hardware, but it’ll be Europe who reaps the most economic value for the cheapest price from their inventions. This is why Europe always appears to be an economic laggard, yet by all metrics it has the best quality of life for the most people anywhere in the world despite having the lowest debt to GDP ratio of any of the world superpowers. Before some say ‘that’s because you don’t spend enough on defence’, I already debunked that in past posts here: Europe has rarely spent less than the US on a PPP adjusted basis, and last few years it is by far the biggest military spender in the world (and if you include Russia in Europe, which most would, then Europe has by far and away always spent more on its military than anywhere else in PPP terms). So, in terms of economic and welfare achievement, Europe’s practice of cheaply reaping from what others sow has served it very well.
How will this affect individual behaviours and mentality?
What will the world be like when your laptop and increasingly your phone locally runs a LLM as powerful or more powerful than the world’s currently most powerful LLM?
You might think what is different to the laptop or phone using a LLM running in a cloud elsewhere and using it over a data connection?, and in some ways you would be right: I’m using a Deepseek running in some cloud elsewhere over a network connection. What’s the difference between that and running it locally?
The first difference is privacy: I wouldn’t ever put anything potentially confidential anywhere near a public internet connection. I definitely wouldn’t put any personal emails or family photos near a public internet connection. Most people won’t care, so maybe this difference only matters to people like me. Still, I’m also an individual, and for me this matters a lot.
The second difference is cost centring: if a cloud runs the LLM, somebody has to pay for that and your average individual is highly adverse to subscriptions when a free of cost substitute is available. So 98% of individuals right now use the free LLM services, and they are generally terrible because otherwise they’d cost real money. If the LLM runs on your device, you take a hit to battery life, but otherwise it’s free of cost. So for your typical individual, from 2028 onwards they’re going to experience an enormous leap in LLM capability, as until then all they’ll be used to is the crappy cheap to operate free LLMs.
The third difference is that most businesses – and a fair few individuals – don’t like to introduce single points of failure to their operations. Most cloud services are seen by many as exceedingly annoying when they go down. And the more you depend on such a service, the more anxious you get if it could disappear/get cut off/drop out. If LLMs run exclusively on hardware you personally own and control, a lot of that anxiety lifts. Now you can lock yourself into this new technology with a certainty you couldn’t have had before. It is for this exact reason why private automobiles are so popular: you aren’t buying transport from A to B, rather you’re buying the guarantee of transport from A to B which public transport only offers in big cities.
The fourth difference is that if they’re truly free of cost and you can run them all day long and all it costs you is electricity, you’re going to use them a LOT more. For everything in fact. Why search the web if your local LLM can do it for you? Why order anything or reply to any message if your local LLM can do it for you? Why think about interacting with your device if your local LLM can do it for you?
And now we’re getting into the interesting stuff: what can a LLM automate away, and what can’t it do i.e. what role is left for humans?
We don’t still know how intelligent LLMs will become before the bubble pops, but I can say this: the LLM knows more about everything than you do, but not more about some specific topics than you do. Accepting on what topics you are weak but being honest about where you genuinely really do understand more than the LLM will be the key to your success going forth.
LLMs genuinely can be a force multiplier if you use them where their strengths lie, and combine that with your strengths. But they also hallucinate and are currently lousy at direction and strategy, so that’s where I would expect the value of humans to remain. In other words, I think politicians are going to have some of the best job security going forth, because their whole purpose is to set unpopular directions for everybody else.
How will this affect individual employment?
The future world of human employment I suspect is (a) those physical jobs which can’t economically be replaced by a robot controlled by a LLM and (b) those jobs where a human’s deep understanding of a niche topic of value cannot be surpassed by any LLM, or where decisions must be taken which involve long term direction and strategy. For everything else, I expect LLMs to gradually replace all before them.
Speaking of LLM controlled robots, I was quite surprised to discover that they only melded a LLM with a humanoid robot last June, so we’ve got a few years to go before humanoid robots start taking human physical jobs. But not as many years as you might think!
Much also to my surprise, it turns out that nobody was mass producing
non-toy humanoid robots until only November last year! Absolutely before
then as now you can buy toy humanoid robots,
these will dance for you
and do kung fu etc, but they’re absolutely useless for getting any real
work done as they (a) can’t lift enough weight reliably and safely and (b)
they don’t have the sensors for fine dexterity manipulation in unfamiliar situations.
And absolutely before now there were intelligent mass produced
industrial robots – any modern factory is stuffed with them – but none
were humanoid until last November.
The first mass produced industrial humanoid robot was the Ubtech Walker S2 which you can see to the right, and it went on sale in November 2025 and has probably sold about three thousand units. It costs about €150k ex VAT, it can carry up to 15 kg and you get about 2.5 hours per battery charge, though it can swap out its battery at a battery recharge station on its own so it can work continuously without a break. Its intended use is within pristine environments where fairly fixed programming works well e.g. walk over there, pick up one of X, rotate it until it has the right orientation, walk back here, put it into the right component box. In other words, just like any other industrial robot, but this one is capable of adapting to different locations within the same factory.
Next up is the Boston Dynamics Atlas which entered mass production in January 2026, and probably about two thousand units have been sold so far. It costs about €200k ex VAT, it can carry up to 30 kg including an impressive 20 kg if on one arm, and you get about two hours per battery charge. It seems a bit more intelligent than the Walker S2, but not by much: it is also intended for fixed, repetitive, work in a pristine environment like a factory floor. This robot is undoubtedly a lot more impressive in the build quality sense than the Walker S2, but it does cost a third more, and also Boston Dynamics only put it into mass production now after decades of development because they had to due to Chinese competition – not because it was finished or particularly compelling or priced well. It also is less interesting because all its production for the next two years is already sold to Hyundai, so nobody else will be able to buy one for several years more yet.
Last April, a much more interesting industrial humanoid robot went into mass production: the Figure 03. Here is a youtube live stream of it unpacking parcels in a mail office, ensuring that the address label points downwards for scanning:
Firstly, the handling of irregularly sized, sometimes squishy, items is FAR harder than the regularly sized boxes with grab handles that the previous two robots can handle. The Figure 03 can also climb stairs by itself – albeit slower than a very elderly person – but it does get there. It currently costs about €100k ex VAT, it can carry up to 20 kg and you get a very good five hours of battery life, but at the cost of it being 40% slower at movement e.g. it walks slower, moves slower etc. They have sold maybe four thousand of these by now. It comes with a bundled LLM running locally which isn’t particularly good – nowhere near even Deepseek v4 Flash in fluid conversation – but if you tell it to go wash the clothes in the washing basket it’ll go fetch the basket and take it to the washing machine, very slowly pick each item out and put it into the machine, then very slowly pour in detergent and set the washing machine running. Ultimately, apart from the thing getting in the way a lot due to its lethargy, it is a far more interesting humanoid robot – at least for the very wealthy, not least due to its cost, but also because you would really need a home with large open spaces so you can easily get around the robot while it very slowly does things. Just to be clear: the Figure 03 can run and jog as fast as a human, but it absolutely horses through its battery if it does, plus it gets hot – very hot! Still, maybe future firmware revisions could let it exchange battery for speed for short bursts so it isn’t annoying and doesn’t get in the way, but otherwise conserve battery life.
If I am being honest though, I suspect the Figure 04 is the one to wait for, as the Figure 03 feels like it has too many design and hardware compromises, and it is still too expensive for what you get on the software side. If it’s the most impressive in mass production right now, what screams out loudly is just how immature and unfinished its software story is. It’ll be years, at best, before that can be remedied.
In case you’re wondering what about all the other mass produced industrial humanoid robots, that’s it: everything else isn’t actually in mass production. In particular, Tesla’s very long advertised robot is nowhere to be seen: we don’t know its specs, its price, or anything else about it, and given Elon Musk’s long history of made up claims about autonomous driving, I wouldn’t be optimistic that his robot will have good autonomy for at least several years after launch. Ultimately this is because it’s one hard thing to build the hardware for an affordable price, it’s another hard thing to create compelling software for that hardware platform. As an example, Meta solved building affordable VR headset hardware, but they did not solve building a compelling software ecosystem for it, so the whole thing has gone off to die and it’s only a matter of time before that entire ecosystem is abandoned. Similarly, Tesla’s fully autonomous driving will likely never get solved well enough to be allowed by regulators at a price consumers will pay.
As much as Figure 03 is impressive, it still requires a pristine environment i.e. you can’t be taking it onto a building site. Even if a robot could navigate well such an irregular environment, and it coped well with getting mud and sand into its joints, it would almost certainly move too slowly for many tasks on a building site AND generally annoy the human construction workers by getting in the way.
Currently a construction worker might cost about €100k to the employer, so maybe for €100k a construction site robot might be worth the expense if it only did things like fetch concrete blocks so the blocklayers could keep working without pause. But as each block weighs 30 kg, it would need to be able to move wheelbarrows of them over scaffolding, which is far beyond the capabilities of any current or near future humanoid robot. You’d also need several of them as they’d go much slower than humans, and because they’d run out of batteries after a few hours you’d either need a quick battery swap facility or even more robots. And finally most construction sites don’t have electricity apart from a generator, so charging robots at a site would be very unattractive. So, for certainly the next decade, I think construction workers can rest in peace that they will not get made unemployed by humanoid robots.
For human jobs doing physical labour in pristine environments though, the next ten years looks like increasing levels of human jobs being displaced. If they can get the cost of these robots down to €25k, a lot of minimum wage jobs like stacking shelves or packing online orders look inevitably gone forever, as the minimum cost to the employer of a human (minimum wage is about €32k) makes the robot look cheaper. If somebody successfully cracks deep cleaning by robot, that’s all your cleaning staff gone too. Jobs like fast food kitchen work is at threat, even if the delivery driver is not – that’s a lot of your young person entry level jobs disappearing forever there.
The most recent (2024) ESRI report lists the largest number of minimum wage jobs being in these sectors:
- Kitchen helpers (14%)
- Shop sales assistants (10%)
- Bartenders (7%)
- Caretakers (6%)
- Waiters (6%)
- Home based personal care (3%)
- Housekeepers (2%)
- Receptionists (2%)
About ten percent of the Irish workforce earns near minimum wage, and humanoid robots could take over most if not all those eight sectors if they get cheap enough.
Food for thought indeed! This displacement of humans from their jobs by AI might have impacted IT first, but I am extremely sure it shall be coming for entire sectors of knowledge worker and pristine environment manual labour next.
| Go to previous entry | Go back to the archive index | Go back to the latest entries |