Niall’s virtual diary archives – Friday 7 August 2026

by . Last updated .

Word count: 18894. Estimated reading time: 89 minutes.
Summary:
A diary entry is provided. The release and performance of Deepseek v4 Flash is discussed. Comparisons are made with other LLMs through coding and summarisation tests. Predictions regarding AI evolution, future Apple hardware, and the impact of humanoid robots on employment are detailed. The diffusion of technology within Europe is also analysed.
Friday 7 August 2026:
00:01.
Word count: 18894. Estimated reading time: 89 minutes.
Summary:
A diary entry is provided. The release and performance of Deepseek v4 Flash is discussed. Comparisons are made with other LLMs through coding and summarisation tests. Predictions regarding AI evolution, future Apple hardware, and the impact of humanoid robots on employment are detailed. The diffusion of technology within Europe is also analysed.
I hadn’t expected to write another diary entry in here during August, however last week Deepseek v4 Flash 0731 was released and for reasons as you’ll see shortly, I have spent most days this past week using it all day and sometimes all night long. We may finally have a Step 3.7 Flash LLM replacement!

Before we get into that though, yesterday my children finished painting the west wall white, thus concluding successfully the painting of around one hundred square metres of exterior wall. I think they did really great given their ages:

To complete a job like this over multiple days, it requires a focus and self control and willingness to see things through to when they are complete which I find lacking in most eighteen year olds, never mind much younger again. Well done Clara, Henry and Julia!

Cheap open weights AI leaps forward yet again!

You may remember that I was initially keen on Qwen3 Coder Next, it was rather slow on my ancient hardware but it did work. However I found myself thereafter mostly using Step 3.5 Flash rented from OpenRouter as it was surprisingly good at coding and agentic work, and it looks like I was early compared to most to realise this – however, then Step 3.7 Flash dropped, and it was better in every way however also twice as expensive for new input BUT now they had prompt caching implemented. Step 3.7 also emitted far fewer thinking tokens than Step 3.5, so all in all the actual cost paid dropped by about half, and I’ve found myself using Step 3.7 Flash for pretty much everything since its release as it had the best ‘bang for the buck’ from my testing i.e. Pareto optimum, and to be specific:

  • It is not the most capable model by any means.
  • It makes many mistakes in the code it writes.
  • It can take multiple attempts to perform an edit or call a tool successfully.
  • BUT if you apply multiple rounds of it checking its work it does catch 98% of the bugs and bad logic it writes and fixes them correctly.
  • It is sufficiently cheap that I’ve spent a total of US$14 ever on it, and that’s despite it horsing through 160 million tokens …
  • From my testing on my own actual use cases, it was the optimal cost-benefit choice of LLM for all tasks where the data it processed is public (I use a local small Gemma 4 LLM for anything processing data which isn’t already on the public internet)
  • Things I really like about Step 3.7 Flash: it follows instructions well, it’s very hard to jail break it out of its system prompt, if you order it to be biased or non biased in its system prompt it does as it is told, and a 196b model is feasibly likely to be runnable on consumer hardware arriving soon, so it’s worth investing into mastering this class of LLMs as your daily driver.

Amazingly, it was only five months ago that Qwen3 Coder Next (Q3CN) landed; and just three months since Step 3.7 Flash landed. Now we have the final release of Deepseek v4 Flash, and here are those LLMs compared so you can see why everybody including me is so excited by this particular LLM release and why social media (or at least my view of it) has been jammed with Deepseek v4 Flash 0731 posts for the past week:

Qwen3 Coder NextStep 3.5 FlashStep 3.7 FlashDeepseek v4 Flash 0731 Claude Fable 5
Released:Feb 2026Feb 2026May 2026August 2026June 2026
MoE weights (total-active):80b-a3b196b-a11b196b-a11b284b-a13b6t-a400b
Max input context:262k262k262k1M1M
Typical Openrouter input cost after prompt caching:$0.103/M$0.100/M$0.053/M$0.030/M$3.36/M
Artificial Analysis Intelligence Index:21.12630.349.959.9
Artificial Analysis Analysis Index:36.2?39.669.176.5
Artificial Analysis Agentic Index:8.8?21.545.752.8
AA Omniscience Accuracy:15.8%23.9%25.4%37.2%61.4%
AA Omniscience Non-Hallucination Rate:9.1%14.8%15.6%15.6%45.1%
SciCode:32.3%40.4%40.0%49.9%60.2%

For comparison, I placed in the final column the current best performing LLM anywhere which is Claude Fable 5. It is 112x times more expensive than Deepseek v4 Flash 0731! Until now the Pareto optimum Step 3.7 performed about half as well as state of the art – now you have something 80% as capable and for nearly half the cost of the previous Pareto optimum.

Deepseek v4 Flash 0731 is a model which at 284 billion parameters is still within the realm of near-future consumer hardware: by 2028, as you’ll see later on in this diary entry, your standard new Apple Macbook Pro from 2028 onwards is expected will include similar compute and memory bandwidth to a 2017-era nVidia Volta AI accelerator board. That should run a model like Deepseek v4 Flash well at around one hundred tokens generated per second, and maybe four thousand tokens parsed per second. That’s a good bit faster than my rented edition has been, so that’s more than fast enough for serious usage.

Between now and then, and especially as the price per token has just halved again with Deepseek v4 Flash 0731 while the capabilities took another leap forwards, it makes the most sense to rent. After this AI investment bubble bursts, I fully expect prices for renting LLMs to crash spectacularly, almost to the point of free … which may make buying local LLM capable hardware a tough ask especially if somebody invents an end-to-end cryptographically secure LLM execution engine, which I’m sure is just a matter of time. However, if your new Apple Macbook Pro just comes bundled in for no extra cost the capability to run local LLMs in the hundred of billion parameter class, then that’s a different matter again.

So, this new LLM aces the synthetic benchmarks. How does it perform in real life?

Test 1: Summarising a 19th century will

Replicating this test from last January, I had it summarise the raw transcribed will of my ancestor Walter Irvine which is early 19th century legal English full of transcription errors. Even the ancient llama 3.1 8b doesn’t do terribly at this, but at the time Claude Sonnet 4.5 did a great job. How do Deepseek v4 Flash 0731 and Step 3.7 Flash perform?

Cloud full fat Claude Sonnet 4.5 400b? 8 bit?Step 3.7 Flash 198b 8 bitDeepseek v4 Flash 0731 284b 4 bit
(from the previous post, it was said at the time it was nearly instant) Tokens per second ~220.4 tokens/s
Token count 5987 tokens
Cost $0.0091834875
Duration 27.2s
Tokens per second ~178.6 tokens/s
Token count 4347 tokens
Cost $0.00182941902
Duration 24.3s

This is a comprehensive probate document from Edinburgh dated November 17, 1825, concerning the estate of Walter Irvine, Esquire, who died on January 7, 1824, at Luddington House, Egham, Surrey. He held properties in the Island of Tobago (West Indies), Luddington House in Surrey, and Denino in Fife, Scotland.

Scottish Estate Inventory (Total: £5,258 6s 2d)

I. Rents and Arrears Due (£4,569 13s 2½d):

  • Robert Gray: £654 9s 6d
  • James Quirie: £455 4s
  • Thomas Wilson: £440 14s 4d
  • Alcy Gosman: £982 10s 3d
  • John Reckie: £763 5s 4d
  • John Hortune: £342 15s 2d
  • Balkaithly Park rents
  • D Gosman feuduty: £8 18s 4d

II. Bond from John M. Harlaw of Kirkton: £245 5s 1d plus interest to January 1824 (Total: £262 15s 1½d)

III. Bond by Road Trustees (Anstruther-Crail to St. Andrews): £150 principal plus interest (Total: £162 7s 4½d)

IV. Bond by Lundenfoot Road Trustees: £50 with compounded interest from 1821-1824 (Total: £55 1s 2d)

V. Proceeds from sale of wood: £33 10s

VI. Balance due from Walter Dickson WS (Writer to the Signet): £174 16s 11d plus interest (Total: £174 18s 11d)

Will Provisions (dated October 18, 1823)

Primary Beneficiaries:

Wife - Catherine Irvine (née Gordon):

  • Annual annuity of £1,500 for life (one-third paid from Scottish estates, two-thirds from daughters' inheritances)
  • Lump sum of £1,000 (£500 within 3 months, £500 within 6 months of death)
  • Lifetime use of all carriages, horses, household furniture, plate, linen, china, books, and consumable wines/liquors in Great Britain
  • Right to occupy Luddington House for life, or receive rents if she chooses not to occupy (£400 annually if the property is sold)
  • This provision replaces her marriage settlement rights and any dower claims

Three Legitimate Daughters (each receiving £35,000):

  1. Elizabeth Douglas (married to William Robert Keith Douglas):

    • Subject to marriage settlement of November 23, 1821
    • Receives one-third value of all Walter Irvine's estates, calculated primarily from the Fife properties
    • Her portion held for her separate use, independent of her husband
    • She has full power to dispose of it by deed or will despite coverture (marriage)
  2. Christina Charles and Catherine Grace (unmarried):

    • Each receives £35,000 legacy
    • Held in trust, with income for their separate use during their lives
    • Cannot be anticipated or attached by creditors or future husbands
    • If married, their surviving husbands receive life interest (only if marriage approved by trustees and if first husband)
    • Upon their deaths, portions pass to their children; if no children survive, portions revert to Elizabeth or are divided among sisters

Natural Daughter - Betsey (of London):

  • Trust fund of £3,333 6s 8d in 3% consolidated annuities
  • Receives income for life for her separate use
  • After her death, capital passes to her children
  • She and any husband can jointly appoint how funds are distributed to children
  • Also receives £25 for mourning

Other Significant Bequests:

Ann Stewart (sister-in-law and widow of Stewart):

  • £50 for mourning
  • Original annuity of £100 increased to £200 per year for life
  • Paid for her separate use, independent of any husband
  • Funded by £6,668 13s 4d in 3% consolidated annuities

Brother's Estate:

  • Forgives all debts owed by the estate of his late brother Christopher William Irvine

Nephews:

  • Christopher Irvine (of Tobago): £4,000 payable one year after death with 4% interest thereafter
  • Walter Irvine: £3,000 under same terms

Nieces (Clara and Marian Irvine - daughters of Christopher William):

  • Each receives annuity of £100 per year for life
  • Each funded by £3,333 6s 8d in 3% consolidated annuities

Jane Armstrong (formerly Jane Currie, wife of Francis Armstrong):

  • £100 lump sum within 3 months for her separate use
  • Annuity of £50 per year for life, paid for her separate use
  • Funded by £1,666 13s 4d in 3% consolidated annuities

Mary (reputed daughter of a negro woman named Linda in Tobago):

  • £500 sterling within 6 months of death

Other Named Beneficiaries:

  • Alexander Gordon (brother-in-law and executor): £500 for his trouble executing the will (£100 within 6 months, £400 in annual installments)
  • John Macfarlane, Esquire: £50
  • Grace (wife of David Maitland), Daniel Gordon, Euphemia Gordon, and Elizabeth Gordon: £200 each plus £25 each for mourning
  • Susannah Ann Hamilton (daughter of executor John Hamilton): £100 (added by codicil)

Continuing Annuities:

  • Mrs. McIntyre: £5 annually during her life
  • Two Miss Irvines: 3 guineas each annually during their lives
  • All funded by appropriate 3% consolidated annuities

Trust Structure and Powers:

Executors and Trustees:

  1. Catherine Irvine (wife)
  2. William Robert Keith Douglas (son-in-law)
  3. Alexander Gordon (brother-in-law, of Old Broad Street, London)
  4. John Hamilton (of Kempshot Park, Hampshire)
  5. Christopher Irvine (nephew, of Tobago)
  6. Thomas Bruce (of George Muir, County of Fife)

Powers Granted to Trustees:

  • Full discretion to manage and invest trust funds in British government stocks, real securities in England or Scotland
  • Authority to purchase lands in fee simple or long leases (100+ years remaining)
  • Power to convert land back to money through sales
  • Can make leases up to 21 years
  • May advance up to one-fourth of any child's portion for education, profession, military commission, or marriage
  • Can use income from children's portions for their maintenance and education during minority
  • William Robert Keith Douglas given principal management of West Indies estates
  • Alexander Gordon entitled to usual solicitor's charges despite being a trustee
  • Power to appoint new trustees upon death or resignation

Specific Provisions for Children's Trusts:

  • Income held for separate use of each daughter, protected from husbands' debts
  • Children of daughters inherit at age 21 (sons) or age 21/marriage (daughters)
  • Compound interest applied to accumulate portions until vesting
  • Complex survivorship provisions if daughters die without surviving children

West Indies Property:

John Leith Dispute:

  • Walter Irvine had claims on property of late John Leith of Tobago
  • Instructs executors to settle these claims quickly
  • Prefers cash payment over taking possession of Old and New Grange estates in Tobago

Plantation Management:

  • Douglas given control of West Indies estates, consignment of produce, and provision of supplies

Scottish Property Requirements:

Election Doctrine:

  • All beneficiaries must execute proper Scottish deeds within 6 months of death (or majority for younger daughter)
  • Those refusing forfeit benefits under the will to compensate disappointed parties
  • Necessary to enable trustees to carry out the will's provisions in Scotland

Residuary Estate:

After all debts, legacies, and annuities paid, the residue divided equally among three legitimate daughters as tenants in common:

  • Elizabeth Douglas
  • Christina Charles
  • Catherine Grace

Elizabeth's share for her separate use with power to dispose by deed or will despite coverture.

Probate Details:

Proved: January 28, 1824, in Prerogative Court of Canterbury

Initial Grant of Administration to:

  • Catherine Irvine (relict/widow)
  • William Robert Keith Douglas
  • Alexander Gordon

Reserved Power for: John Hamilton, Christopher Irvine, and Thomas Bruce to apply later

Witnesses to Will:

  • Colvell Wyatt Edgell (of Milton Place, Egham, Surrey)
  • Robert Logan (Egham Lodge)
  • William Adams (servant to Mr. Irvine)

Codicil: Dated same day (October 18, 1823), adding £100 bequest to Susannah Ann Hamilton

Scottish Confirmation: Presented November 17, 1825, before George Carshin Jr., Commissioner of Edinburgh, with deposition by W.R.K. Douglas confirming the inventory and will details.

The document demonstrates extensive wealth with significant properties in Scotland, England, and the West Indies, careful provision for family members including an illegitimate daughter, and sophisticated trust arrangements typical of wealthy estates in the early 19th century.

Document Date: 17 November 1825 (Edinburgh)
Testator: Walter Irvine, Esquire (late of Tobago, Luddington House, Surrey, and Denino, Fife)
Date of Death: 7 January 1824
Will Date: 18 October 1823 (with a codicil of the same date)
Probate: Granted 28 January 1824 by the Prerogative Court of Canterbury


1. Inventory of Personal Estate (Scotland)

Presented by executor William Robert Keith Douglas (WS) and signed by him and George Carfrae Jr. (Commissioner). The total value of the Scottish estate is £5,258 / 6 / 2 (£5,258.31).

No. Debtor / Item Amount (£)
I. Rents and arrears (5 debtors + 3 items) £4,569 / 13 / 2½
II. Balance due by John Hay Harlaw (principal + interest at 5%) £262 / 15 / 1½
III. Bond from Trustees of Anstruther & Crail to St. Andrews road (principal + interest) £162 / 7 / 4½
IV. Sum lent to Lauderfoot road Trustees (principal + varying interest) £55 / 1 / 2
V. Proceeds of Roupes of wood £33 / 10 / 0
VI. Balance due by Walter Dickson WS on account current (principal + interest) £174 / 18 / 11

The testator also owned property in England; letters of administration for that part were already taken out by the executors.


2. Deposition (25 October 1825)

Executor W. R. K. Douglas swore that:

  • Walter Irvine died 7 January 1824.
  • He, his widow Catherine Gordon or Irvine, Alexander Gordon, John Hamilton, Christopher Irvine, and Thomas Bruce entered upon possession as executors under the will proved 28 January 1824.
  • He knew of no other settlement or writing relative to the estate except the will and the marriage contract of Walter Irvine and Catherine (date blank).
  • The inventory together with the English property constitutes a full and complete inventory of the personal estate, so far as known to him.
  • The Scottish estate is valued between £5,000 and £6,000.
  • He does not require confirmation to be expedited.

3. Probate

  • Proved 28 January 1824 before the Prerogative Court of Canterbury.
  • Grant to three executors: Catherine Irvine (widow), William Robert Keith Douglas, and Alexander Gordon (all duly sworn).
  • Power reserved to grant similar administration to the other named executors (John Hamilton, Christopher Irvine, Thomas Bruce) upon application.
  • Certificate dated 12 November 1825 signed by W. R. K. Douglas confirming the probate of the will.

4. Last Will and Testament – Detailed Summary

4.1 General Directions

  • Revokes all earlier wills and codicils.
  • All just debts, funeral expenses, and legacies to be paid from personal estate; if insufficient, real estate is charged.
  • All payments to be made with convenient speed unless otherwise directed.

4.2 Provision for Wife – Catherine Irvine

  • Annuity: £1,500 per year (half‑yearly payments on 25 June and 25 December), charged on all estates.
    • Scottish estates settled for daughter Elizabeth bear one‑third of this annuity.
    • Each surviving daughter (or her heirs) must contribute a rateable third of the remaining two‑thirds.
  • Residence: Catherine (and any unmarried daughters) may occupy Luddington House during her lifetime; if she ceases to occupy, the rents are paid to her. She may demise the house for up to three years after her death. If sold during her life, she receives £400 per year in lieu.
  • Cash bequest: £1,000 (half within 3 months, half within 6 months).
  • Household goods: Use for life of all carriages, carriage horses, furniture, plate, linen, china, books in Great Britain, plus wines and liquors consumed in her family.
  • This provision is in lieu of her marriage settlement and bars all dower/thirds rights in England and Scotland.

4.3 Remission of Debt

  • Remits all sums due from the estate of his late brother Christopher William Irvine for the benefit of those liable.

4.4 Bequests to Individuals

Beneficiary Bequest Funding / Notes
Ann Stewart (sister‑in‑law) £50 for mourning; ratification of existing bond for £100 annuity + additional £100 per year (total £200/year) Fund: £6,668 / 13 / 4 3% Consolidated annuities
John Macfarlane £50
Grace Maitland, Daniel Gordon, Euphemia Gordon, Elizabeth Gordon £200 each + £25 each for mourning
Alexander Gordon (brother‑in‑law) £500 for trouble as trustee ( £100 within 6 months, £400 in equal annual instalments of £100)
Jane Armstrong (formerly Jane Currie) £100 (within 3 months, separate use) + £50 per year annuity for life Fund: £1,666 / 13 / 4 3% Consolidated annuities
Clara Irvine & Marian Irvine (nieces, daughters of brother Christopher William) £100 per year each for life Each funded by £3,333 / 6 / 8 3% Consolidated annuities
Mary (reputed daughter of negro woman Linda in Tobago) £500 within 6 months
Mrs. McIntyre £5 per year for life (stamp‑free) Funded by 3% Consolidated annuities
Two Miss Irvines 3 guineas (£3 / 3 / 0) per year for life (stamp‑free) Funded by 3% Consolidated annuities
Christopher Irvine (nephew) £3,000 (payable after 1 year with 4% interest)
Walter Irvine (nephew) £3,000 (payable after 1 year with 4% interest)

4.5 Trust for Natural Daughter Betsy

  • Capital: £3,333 / 6 / 8 in 3% Consolidated annuities.
  • Income: Paid to Betsy for life, for her sole and separate use, free from any husband’s control, without power to anticipate.
  • Remainder: After Betsy’s death, the fund is held in trust for her children/issue as she (and any husband) may jointly appoint; failing appointment, for her children equally (sons at 21, daughters at 21 or marriage). If all such children fail under age without issue, the fund falls into the residue.

4.6 Main Trusts for Daughters Christina Charles and Catherine Grace

  • Legacy: Each receives a legacy of £35,000, to be answered by purchasing 3% Consolidated annuities at the highest market price on the day of death (or last open day). The annuities carry the first dividend after death.
  • Trust: During each daughter’s life, the income is for her sole and separate use, not subject to anticipation or her husband’s control.
  • Surviving husband: If a husband survives, he is entitled to the income for life only if he was the first husband and married with the written consent of the trustees (or at least three of them).
  • Remainder: After death of daughter and any surviving husband, the capital is held for her children/issue as she appoints by deed or will; in default, for her children equally (sons at 21, daughters at 21 or marriage). Further provisions deal with failure of issue (gift over to the other daughter or to the married daughter Elizabeth Douglas).
  • Powers: Trustees may apply income for maintenance/education of minor grandchildren; may advance up to one‑quarter of a child’s share for profession, army commission, university, etc. (treated as a deduction). Undisposed income is to be compounded and added to principal.

4.7 Investment and Management Powers of Trustees

  • May invest in British stocks/funds or real securities in England, Scotland, or Wales.
  • With consent of the life tenant (or after death, at trustees’ discretion) may vary securities.
  • Upon request of the life tenant, may purchase land (fee simple or term with ≥100 years unexpired); may later sell and reconvert to money; may lease lands for up to 21 years; apply rents as if they were interest.
  • The fund may be held on the same trusts as personal estate.

4.8 Devise of Real Estate

  • All freehold, copyhold, leasehold manors, plantations, negroes, slaves, lands, etc., in England, Scotland, West Indies, and elsewhere, plus estates held as mortgagee/trustee, are devised to the trustees (Catherine Irvine, W. R. K. Douglas, Alexander Gordon, John Hamilton, Christopher Irvine, Thomas Bruce) according to their nature, on the trusts declared in the will.

4.9 Executors and Guardians

  • The same six persons are appointed executors in trust and guardians of the youngest daughter during her minority.

4.10 Election Regarding Scottish Estates

  • All three daughters (and other beneficiaries) must, within six calendar months after death (or for the youngest daughter, within six months after coming of age), execute a proper Scotch form deed of disposition of all Scottish real/heritable estates to enable the trustees to carry out the will.
  • Failure to do so binds them, under the doctrine of Election (as observed by English Courts of Equity), to compensate those disappointed by their neglect.

4.11 Covenant for Daughter Elizabeth’s Marriage Settlement

  • The testator had covenanted (in a settlement dated ~23 November 1821) to settle one‑third of all his real and personal estate (worldwide) upon Elizabeth and her family.
  • That covenant is subject to his lifetime power to dispose of property by act or will, and to charge it with annuities/legacies.
  • The trustees (Catherine, Alexander, John, Christopher, Thomas) have sole discretion to arrange which part of the West Indies estates shall be appropriated to make up that third. Their certificate is binding.
  • The third is computed after deducting all monetary legacies and annuity funds (except those created solely for life annuities). This third is the primary fund to pay one‑third of the wife’s annuity, exonerating Elizabeth’s share.
  • Trustees are empowered to cause valuations and to complete the settlement accordingly.

4.12 Residuary Estate

  • After payment of debts, legacies, and fulfilling the marriage settlement, all remaining real and personal estate (including funds appropriated for annuities) is held in trust for the three daughters (Elizabeth Douglas, Christina Charles, Catherine Grace) equally as tenants in common (not joint tenants).
  • Elizabeth’s share is for her sole and separate use; she may dispose of it by deed or will as if a feme sole, regardless of coverture.

4.13 Receipts and Discharges

  • Receipts in writing by any trustee (or their survivors, heirs, agents) for money payable under the will shall effectually discharge the payer from seeing to application or being accountable for misapplication.

4.14 Powers of Executors

  • Executors may pay debts on any evidence they think proper; accept security; compromise debts owed to the estate; allow time for payment.
  • They are directed to accept settlement of claims on John Leith’s Tobago property in preference to taking the Old and New Grange estates in Tobago.

4.15 Appointment of New Trustees

  • If any trustee dies, is discharged, or refuses to act, the surviving/acting trustees may appoint new trustees by deed under seal, attested by two witnesses.
  • The trust estates shall be conveyed to the new trustees, who shall have the same powers as original trustees.

4.16 Indemnity of Trustees

  • Trustees may deduct and retain from trust monies all costs, charges, and expenses incurred in execution of the trusts (approved by the majority of acting trustees).
  • Alexander Gordon, although a trustee, may charge the usual solicitor’s fees for business transacted by him.
  • Trustees are only chargeable for their own receipts and wilful defaults; not liable for loss without wilful default.

4.17 Management Powers

  • W. R. K. Douglas is entrusted with principal direction and management of the West Indies estates (consignment of produce, supply of plantations).
  • Trustees may abate rents, make leases up to 21 years of untenanted or soon‑expiring trust property.

4.18 Execution

  • Signed and sealed by Walter Irvine on 18 October 1823; witnessed by Colgell Wyatt Edgell, Robert Logan, and William Adams (servant to Mr. Irvine).

5. Codicil (18 October 1823)

  • Adds a bequest of £100 to Susanna Ann Hamilton, daughter of John Hamilton Esquire.

6. Affidavits and Probate Details

Handwriting Affidavits (for the codicil)

  • William Matthews and Joseph Pennington (both of Old Broad Street, London) swore they knew Walter Irvine’s handwriting and believed the signature on the codicil to be his.
  • Sworn before John Daubeny, Surrogate, on 21 and 22 January 1824.

Attesting Witness Affidavit

  • William Adams (servant) swore he was present at the execution of the will; described the signing, sealing, publishing, and witnessing; confirmed the testator was of sound mind. Sworn 21 January 1824.

Grant of Probate

  • Granted to Catherine Irvine, W. R. K. Douglas, and Alexander Gordon on 28 January 1824.
  • Power reserved to the other executors (John Hamilton, Christopher Irvine, Thomas Bruce) to apply later.

7. Key Figures at a Glance

  • Total Scottish personal estate: £5,258 / 6 / 2
  • Wife’s annuity: £1,500/year
  • Two younger daughters’ legacies: £35,000 each (to be invested in 3% consols)
  • Natural daughter Betsy’s fund: £3,333 / 6 / 8 3% consols
  • Nephews (Christopher & Walter Irvine): £3,000 each
  • Nieces (Clara & Marian): £100/year each (funded by £3,333 / 6 / 8 each)
  • Jane Armstrong: £100 + £50/year annuity
  • Mary (daughter of Linda): £500

This document is a comprehensive and intricate testament disposing of a substantial trans‑Atlantic estate, providing generously for the testator’s wife, daughters (both legitimate and natural), other relatives, and servants, while establishing elaborate trusts to preserve wealth across generations.

This document is a collection of probate records concerning the estate of Walter Irvine, an Esquire who held property in the Island of Tobago (West Indies), Luddington House in Surrey, and Denino in Fife, Scotland. He died at Luddington House on 7 January 1824. The documents date from between late 1823 and late 1825 and comprise an inventory of his Scottish personal estate, a sworn deposition, his last will and testament (dated 18 October 1823) with a codicil, supporting affidavits, and the formal grant of probate from the Prerogative Court of Canterbury.


I. Inventory of the Personal Estate (Scotland)

The inventory was presented in Edinburgh on 17 November 1825 by Mr. Walter Dickson, Writer to the Signet, and made out by William Robert Keith Douglas, Esquire, one of the Lords of the Admiralty and an executor under the will (which was proved in the Prerogative Court of Canterbury in 1824, under the Act 48 Geo III cap 140).

I. Rents and Arrears of Rent due to the deceased

Debtor Amount
Robert Gray £654. 9. 6
James Currie £455. 4. 0
Thomas Wilson £440. 14. 4
Alex Gossman £982. 10. 3
John Reekie £763. 5. 4
John Fortune (Balkello Park rents &c) £342. 15. 2
D. Gossman (feu duty) £8. 18. 4
Total £4569. 13. 2½

II. Balance of debt due by John Hay Harlaw, Esquire of Kirkton

  • Principal (at 4 Aug 1822): £245. 5. 1
  • Interest to 7 Jan 1824 at 5%: £17. 10. 0½
  • Total: £262. 15. 1½

III. Sum in bond by the Trustees of the Anstruther & Crail to St. Andrews road

  • Principal: £150. 0. 0
  • Interest (Whitsun 1822 to 7 Jan 1824, 5%): £12. 7. 4½
  • Total: £162. 7. 4½

IV. Sum lent to the Trustees of the Lauderfoot road

  • Principal: £50. 0. 0
  • Interest at varying rates (5%, 4½%, 4% over successive years): £4. 17. 5
  • Total: £55. 1. 2

V. Proceeds of roupes of wood due at death: £33. 10. 0

VI. Balance due by Walter Dickson WS on account current (31 Dec 1823)

  • Principal: £174. 16. 11
  • Interest at 3%: £0. 2. 0
  • Total: £174. 18. 11

Total value of the deceased's estate in Scotland: £5258. 6. 2

The deceased also died possessed of property in England, for which Letters of Administration had already been taken out by the executors. The inventory was signed by W. R. K. Douglas and George Carfrae Jr.


II. Deposition (Edinburgh, 25 October 1825)

Before George Carfrae Junior, Commissioner appointed by the Commissaries of Edinburgh, William Robert Keith Douglas swore that:

  • Walter Irvine died on 7 January 1824.
  • Douglas, acting with Mrs. Catherine Gordon or Irvine (the widow), Alexander Gordon (Old Broad Street, London), John Hamilton (Kempshott Park, Hants), and Christopher Irvine (of Tobago), entered into possession and management of the estate as executors under the will dated 18 October 1823.
  • The will was proved in the Prerogative Court of Canterbury on 28 January 1824 (the probate was then in London but would be transmitted before the inventory was recorded).
  • He knew of no other settlement or writing disposing of the estate other than the will and an (undated) contract of marriage between the deceased and Catherine Gordon.
  • The inventory (together with the English effects) was a full and complete inventory of the deceased's personal estate.
  • The Scottish estate was valued at over £5,000 but under £6,000 sterling.
  • The deponent did not require Confirmation to be expended for the specified debts/effects.

III. The Last Will and Testament (18 October 1823)

The will revokes all prior wills, codicils, and testamentary dispositions. It directs all just debts, funeral expenses, probate costs, and legacies to be paid promptly out of the personal estate, with the real estate charged for any deficiency.

Provision for the Wife — Catherine Irvine

  • An annuity of £1,500 per annum for life, paid half-yearly on 25 June and 25 December.
  • Charged on all estates; the Scottish estates settled on daughter Elizabeth bear one-third of it, and each of the three surviving daughters bears a rateable third of the remaining two-thirds.
  • The wife and unmarried daughters may occupy Luddington House for life; if she vacates, rents go to her to fund another residence, and trustees may lease the property for terms ending within three years after her death. If sold in her lifetime, she receives £400 a year in lieu.
  • A legacy of £1,000 (paid in two moieties within three and six months).
  • Use of carriages, carriage horses, household furniture, plate, linen, china, books, and part of the wines/liquors for life.
  • This provision is in lieu of her marriage settlement and bars all dower/thirds rights in both England and Scotland.

Remission of Debt

  • All sums due from the estate of his late brother Christopher William Irvine are remitted for the benefit of those liable.

Bequest to Ann Stewart (sister-in-law)

  • £50 for mourning.
  • The existing bond annuity of £100 p.a. is ratified, and an additional £100 p.a. (already being paid) is confirmed for life — a total of £200 p.a.
  • A fund of £6,668 13s 4d of 3% Consolidated annuities is appropriated as the only fund to answer this annuity.

Other Bequests

  • John Macfarlane, Esquire: £50.
  • Sister-in-law Grace (wife of David Maitland), Daniel Gordon, Euphemia Gordon, and Elizabeth Gordon: £200 each plus £25 each for mourning.
  • Alexander Gordon (brother-in-law): £500 for trouble in executing the trusts (£100 within six months, the rest in annual instalments of £100).

Trust for Natural Daughter "Betsy"

  • A trust of £3,333 6s 8d of 3% Consolidated annuities for his reputed daughter Betsy of London, for life, for her sole and separate use, free from any husband's control, without power to anticipate.
  • After her death, the fund goes to her children/issue as she and any husband may jointly appoint (by deed or will); in default, equally among her lawful children at 21 (sons) or 21/marriage with consent (daughters), as tenants in common, with survivorship.
  • If she has no children, the fund falls into residue.
  • £25 extra for mourning for Betsy.

Bequest to Jane Armstrong (formerly Jane Currie)

  • A lump sum of £100 (paid within three months) for her sole and separate use, free from her husband Francis Armstrong.
  • An annuity of £50 p.a. for life, half-yearly on 5 July and 5 January, for her separate use, backed by a fund of £1,666 13s 4d of 3% annuities as the only fund.

Bequests to Nieces Clara and Marian Irvine

  • Each of the two daughters of his late brother Christopher William Irvine receives an annuity of £100 p.a. for life (half-yearly on 5 July and 5 January), each backed by a fund of £3,333 6s 8d of 3% annuities as the only fund.
  • The wife retains a general lien on all residuary estates for her annuity notwithstanding the appropriations.

Bequest to Mary (reputed daughter of enslaved woman Linda)

  • £500 (paid within six months) to Mary, the reputed daughter of the negro woman Linda living in Tobago.

Annuities to Others

  • The annual £5 to Mrs. McIntyre and the annual 3 guineas to the two Miss Irvines are continued for life, free of stamp duties, each backed by appropriated 3% annuities as the only fund.

Bequests to Nephews

  • Christopher Irvine and Walter Irvine each receive £3,000 (paid one year after death, with 4% interest from that point until payment).

Main Legacies for Daughters Christina and Catherine

  • Two separate legacies, each of £35,000, valued by 3% Consolidated annuities at the highest market price on the day of death (or the last open day if shut), carrying dividends from the first dividend day after death.
  • One legacy is held in trust for daughter Christina Charles, the other for Catherine Grace, on protective trusts: income for each daughter's sole and separate use during life, without power to anticipate, free from any husband's control.
  • A surviving husband may take income for life (provided he married with the written consent of the trustees).
  • After each daughter and her surviving husband die, the legacy goes to her children/issue as she appoints by deed or will; in default, equally among children as tenants in common (sons at 21, daughters at 21 or marriage), with survivorship provisions.
  • On failure of issue, portions are divided between the other daughters' shares, ultimately accumulating toward daughter Elizabeth's settled portion.

Powers of Trustees (Maintenance and Advancement)

  • Trustees may apply income for minor children's maintenance, education, schooling, clothing, or advancement.
  • They may advance up to one-fourth of a child's vested or expectant share for placing a male into a profession, purchasing a military commission, university education, Inns of Court, or a child's marriage/advancement.
  • Advances are deducted from the child's portion; capitalized income accrues with compound interest and follows the same trusts.

Investment Powers

  • With the life-tenant's consent (and after their death, on trustees' own authority), trustees may invest in British stocks/funds or real securities in England, Scotland, or Wales.
  • They may vary/change securities, and (on request of life tenants) purchase or sell land held in fee simple or with at least 100 years unexpired, lease land (up to 21 years), and hold it on the same trusts as personal estate.

Devise of Real Estate

  • The will devises all freehold, copyhold, and leasehold manors, messuages, farms, plantations, negroes, and slaves, lands and hereditaments in England, Scotland, and the West Indies (and estates held as mortgagee/trustee), and all goods and personal estate, to the six trustees (Catherine Irvine, W. R. K. Douglas, Alexander Gordon, John Hamilton, Christopher Irvine, and Thomas Bruce) on the stated trusts.

Appointment of Executors and Guardians

  • The six named individuals are appointed executors in trust and guardians of the youngest daughter during her minority.

Election Requirement for Scottish Estates

  • The three daughters and all beneficiaries must, within six months (or, for the youngest daughter, within six months of coming of age), execute a Deed in Scotch form disposing of the Scottish real/heritable estates so trustees can carry the will into effect.
  • Anyone refusing is bound, under the doctrine of Election, to give compensation/equivalence out of their benefits.

Covenant Regarding Daughter Elizabeth's Marriage Settlement

  • The will directs performance of the covenant in the marriage settlement of 23 November 1821 (between W. R. K. Douglas, Walter Irvine and Elizabeth, and Robert Bruce of Kennet and John Charles Herries) for settling on Elizabeth one-third of all real and personal property in Great Britain, Ireland, the West Indies, and America.
  • This third is to be made up primarily from the Fife estate (regardless of whether it exceeds one-third in value), supplemented if necessary from West Indian plantations/negroes/slaves.
  • The trustees (or any three) have sole discretion over the valuation and appropriation of West Indian estates; their certificate is binding and conclusive.
  • The settled third is the primary fund to answer one-third of the annuities, in exoneration of Elizabeth, and the trustees are to complete the settlement following valuation.

Residuary Estate Distribution

  • Subject to debts, legacies, and regulations, all real and personal estate (including funds appropriated for annuities) is held in trust for the three daughters — Elizabeth Douglas, Christina Charles, and Catherine Grace — or those surviving, equally as tenants in common (not joint tenants).
  • Elizabeth's share is for her sole and separate use, with full power to dispose of it by deed or will as if a feme sole.
  • Reversions in settled Scottish/West Indian estates and the two £35,000 portions pass to the two unmarried daughters as tenants in common with survivorship on failure of issue.
  • No deduction is made from Elizabeth's one-third in respect of her residuary share.

Receipts, Powers of Executors, New Trustees, Indemnity, Management

  • Receipts: Trustees' written receipts are sufficient discharges.
  • Executors' powers: pay disputed debts as they think proper, accept securities, compromise debts, allow time for payment; and if his claims on the estate of John Leith of Tobago are settled, executors should prefer cash payment over taking the Old and New Grange estates in Tobago.
  • New trustees: may be appointed by deed on death/refusal to act, with full powers.
  • Indemnity: trustees may deduct costs from trust funds; Alexander Gordon (as a solicitor) may make usual solicitor's charges; trustees are liable only for their own acts/wilful defaults, not for losses without wilful default.
  • Management: trustees may entrust W. R. K. Douglas with the principal direction of the West Indian estates, consignment of produce, and supplies; may abate rents and grant leases up to 21 years.

Execution and Attestation

Executed on 18 October 1823 at Luddington House, signed and sealed by Walter Irvine across 19 sheets, in the presence of witnesses:

  • Colgell Wyatt Edgell (Milton Place, Egham, Surrey)
  • Robert Logan (Egham Lodge)
  • William Adams (servant to Mr. Irvine)

IV. Codicil

Dated 18 October 1823, adds a bequest of £100 to Susanna Ann Hamilton, daughter of John Hamilton, Esquire.


V. Affidavits and Grant of Probate

Affidavit (Codicil)

  • William Matthews (Old Broad Street, London) and Joseph Pennington swore they knew Walter Irvine's handwriting and believed the signature "Walter Irvine" on the codicil (annexed, marked "A") to be genuine. Pennington was sworn on 21 January 1824, and Matthews on 22 January 1824, before John Daubeny, Surrogate, and John I. Glennie, Notary Public.

Affidavit (William Adams, subscribing witness)

  • William Adams swore he witnessed the testator sign the will on the nineteenth sheet opposite a wax seal, that the testator sealed, published, and declared it as his last will, and that he (Adams) was of sound and disposing mind, memory, and understanding at the time. Sworn 21 January 1824.

Grant of Probate (28 January 1824)

  • Charles, Archbishop of Canterbury, grants probate in the Prerogative Court of Canterbury before John Daubeny, LL.D., Surrogate of Sir John Nicholl, Knight, Master Keeper or Commissary.
  • The deceased is described as formerly of the Island of Tobago, of Wimpole Street, Middlesex, and late of Luddington House, Egham, Surrey.
  • Administration was granted to Catherine Irvine (the relict), William Robert Keith Douglas, and Alexander Gordon, three named executors, who swore to well and faithfully administer and exhibit a true inventory by 31 July next.
  • Power reserved to later grant administration to the other executors (John Hamilton, Christopher Irvine the nephew, and Thomas Bruce) when they apply.
  • Signed by registrars Nathaniel Goschius, N. C. Cresswell, and George Firmer (Deputy Registers).

VI. Certificate

London, 12 November 1825William Robert Keith Douglas certifies that this is the probate of the will of the late Walter Irvine referred to in his 25 October 1825 deposition on the inventory.


Key Points at a Glance

  • Testator: Walter Irvine, Esquire — properties in Tobago (West Indies), Luddington House (Surrey), and Denino (Fife); also Wimpole Street, Middlesex. Died 7 January 1824.
  • Scottish personal estate value: £5,258 6s 2d (over £5,000, under £6,000); additional English assets subject to Letters of Administration.
  • Executors/Trustees (6): Catherine Irvine (wife), William Robert Keith Douglas, Alexander Gordon, John Hamilton, Christopher Irvine (nephew), Thomas Bruce.
  • Major beneficiaries: wife (annuity of £1,500 p.a.); three daughters — Elizabeth Douglas (marriage settlement third), Christina Charles and Catherine Grace (each £35,000); natural daughter Betsy (life interest trust); two nephews (each £3,000); plus numerous smaller legacies and life annuities to family, relations, and dependants (including a sum to Mary, reputed daughter of the enslaved woman Linda).
  • Notable features: protective trusts to shield women's income from husbands; appropriation funds (3% Consols) dedicated to each annuity; the "doctrine of Election" for Scottish estates; power to purchase military commissions; and a slavery-era estate composed of plantations, negroes, and slaves in the West Indies.

Scroll down to see more ...

The good news is that it was much harder for me to determine the strengths and weaknesses of each of the two models compared to Claude from last year: Claude Sonnet 4.5 definitely still wins by a long shot, but the gap is much smaller than it was for Qwen 3.0 30b. Something which piqued my interest is why the Chinese models are so keen on outputting blow-by-blow structure of the original document, and I wondered if it is an artefact of the Mixture of Experts (MoE) design. So I also tested Gemma 4 31b which is dense and Gemma 4 28b-a4b which is MoE, and indeed the same blow-by-blow structure appears for the latter. I guess that kinda makes sense? Incidentally, Gemma 4 31b did surprisingly poorly on this test, I had assumed it would beat Deepseek v4 Flash as the Gemma models are well known to be better at English language nuance than the Chinese models, and while yes it did very well at picking out the right essential points from the will, it didn’t pick enough of those essential points despite being told to be detailed. Maybe I needed to say ‘very detailed’? Don’t get me wrong, the quality of Gemma 4 31b’s output was good, but it was short and to the point as it were, and too much short and to the point in fact.

Re: our two models, I think Step 3.7 produces a better structured documents – it is keen on tables – and it is more terse than Deepseek v4 which gives too much irrelevant detail, plus it writes English better in my opinion: less fluff, more densely packed. Deepseek v4 on the other hand did cost one fifth the amount which Step 3.7 did, and it’s not that much worse. Still, Step 3.7 wins this test on quality of output, if you exclude Claude.

Claude from last year is much better at synthesising the document together e.g. it groups daughters together, it has realised one is illegitimate, it orders items in a reasonable priority for most human readers, and it has collapsed all the multiple sections from the original into the minimum possible set. The Chinese models, despite getting towards a similar 400 billion parameters of Claude from last year, have a way to go yet, assuming that they’ll ever get there as they have a MoE design.

Test 2: Analyse a code implementation of a specification and implementation plan

Last few weeks I have been working on atomic_wait() for C, which essentially ports the same feature from C++ 20 into the next C standard – though we shall be adding some additional APIs, as we don’t care much for the C++ API. Myself and fellow committee member Jens Gustedt came up with a draft WG14 proposal paper over a number of weeks, then I iterated having Step 3.7 Flash write a detailed implementation plan for a reference implementation using another hand written reference implementation for a separate WG14 proposal as a template. It did struggle a bit with writing out the plan, and I had to hand hold it a fair bit, but we got there.

The single most important part of the plan file is probably this which describes when a proxy atomic must be used which is indexed via an internal hash table, or whether the atomic wait can be passed through to the platform specific API directly:

Backend 1 byte 2 bytes 4 bytes 8 bytes Hash table needed?
Linux (FUTEX_WAIT/FUTEX_WAKE) For 1-2-byte and 8-byte; futex is 32-bit only (int *uaddr, int val)
macOS (UL_COMPARE_AND_WAIT/UL_COMPARE_AND_WAIT64) For 1-2-byte, or sub-native-width types
Windows (WaitOnAddress) Never — all operand sizes bypass
FreeBSD (UMTX_OP_WAIT/UMTX_OP_WAKE) For 1-2-byte, or sub-native-width types; UMTX_OP_WAKE accepts a count parameter directly
pthreads fallback (pthread_cond_wait) Always — no kernel tracker exists

✓ = kernel primitive available; hash table is bypassed. ✗ = no suitable kernel primitive; must use the user-space hash table.

So, the design’s essential points are:

  1. There are multiple implementation backends for each platform specific API.
  2. The public API is able to pass through to the kernel API directly for some or all atomic types depending on backend.
  3. For the remaining types, an internal hash table maps an atomic’s address in memory to its proxy atomic which IS compatible with the kernel API.

I asked Step 3.7 Flash to implement the reference library using the plan and proposal as guides. It replicated over the mildly changed parts of the hand written source template well enough, but really struggled with designing and writing the implementation. It came up with a poor architecture and lots of repeated and verbose code all of which was unnecessary. So I ripped out all of that, and hand wrote an implementation for Mac OS which has compile time replaceable backends for different platforms. I then asked Step 3.7 Flash to implement the backends for (i) Microsoft Windows (ii) Linux (iii) FreeBSD and (iv) generic pthreads using the Mac OS backend as a template.

Here it did better, though it messed up the calling protocol, swapped a few arguments around, got error handling wrong etc. All not great, but then Deepseek v4 Flash 0731 was announced, and I thought it might be more interesting to see how both models would analyse the current implementation. I asked both this prompt:

Perform a comprehensive review of the implementation, comparing it to the plan and proposal. Write out any deviations, problems, bugs or any other issue into docs/deepseek_v4_flash_0731_analysis.md. Be detailed, also be aware that there is likely a race condition in the implementation which is causing lost wakes. Try to identify the cause without debugging.

Step 3.7 Flash took about four minutes to complete its analysis, taking fifteen steps and costing US$0.075 dollars to process 101k tokens and output 23k tokens. Deepseek v4 Flash unfortunately took my ‘try’ literally, and launched gdb as part of its analysis. So I refined the prompt just for it:

Perform a comprehensive review of the implementation, comparing it to the plan and proposal. Write out any deviations, problems, bugs or any other issue into docs/deepseek_v4_flash_0731_analysis.md. Be detailed, also be aware that there is likely a race condition in the implementation which is causing lost wakes. Identify the cause without debugging i.e. do NOT run gdb, do NOT run test programs - inspect and analyse the code ONLY.

… which worked, and it took ten minutes to complete its analysis, taking sixteen steeps and costing US$0.029 dollars to process 81k tokens and output 39k tokens. That’s about 2.6x cheaper, incidentally.

Its analysis was deeper, but it missed some easy issues Step 3.7 caught:

Step 3.7 Flash 198b 8 bitDeepseek v4 Flash 0731 284b 4 bit

Step 3.7 Implementation Analysis

1. Deviations from Plan/Proposal

1.1 Windows 1-byte and 2-byte operations use hash table fallback

Plan: "Windows: always bypass (all sizes 1–8 bytes). The Windows backend never uses the hash table."

Actual: atomic_wait_windows.c.ipp only defines HAVE_WAIT_ON_ADDRESS_32 and HAVE_WAIT_ON_ADDRESS_64. There are no wait_on_address8, wait_on_address16, wake_by_address8, or wake_by_address16 definitions. As a result, atomic_wait_1, atomic_wait_2, atomic_notify_one_1, atomic_notify_one_2, atomic_notify_all_1, and atomic_notify_all_2 on Windows fall through to atomic_wait_generic / atomic_notify_generic and use the hash table proxy.

Files: include/wg14_atomic_waits/detail/impl/atomic_wait_windows.c.ipp:42-74

1.2 macOS atomic_notify_all return value is platform-fixed to INT_MAX - 1

Plan: "Returns a positive number (possibly 1 + N where the extra info depends on platform)."

Actual: Both wake_by_address32 and wake_by_address64 on macOS return INT_MAX - 1 when max_threads_to_wake != 1. __ulock_wake does not return a count of woken threads, so the implementation cannot know how many were woken. Returning a fixed large value is not consistent with the plan's "possibly one plus the number of threads woken" semantics.

Files: include/wg14_atomic_waits/detail/impl/atomic_wait_macos.c.ipp:106, include/wg14_atomic_waits/detail/impl/atomic_wait_macos.c.ipp:123

1.3 Windows atomic_notify_all calls WakeByAddressSingle instead of WakeByAddressAll

Plan: "On Windows, call WakeByAddressSingle / WakeByAddressAll directly without any wake_waiters helper."

Actual: wake_by_address32 and wake_by_address64 always call WakeByAddressSingle, regardless of max_threads_to_wake. WakeByAddressAll is never called.

Files: include/wg14_atomic_waits/detail/impl/atomic_wait_windows.c.ipp:100, include/wg14_atomic_waits/detail/impl/atomic_wait_windows.c.ipp:117

1.4 atomic_wait_common.ipp.ipp lacks warning suppression for volatile-to-void* cast

Plan: "Test under -Wextra to suppress qualifier-discard warnings."

Actual: The cast void *const key = (void *) object; on line 173 discards both volatile and const qualifiers. With -Wextra, this produces a warning. No #pragma or explicit cast-to-uintptr_t-then-to-void* is used to suppress it.

Files: include/wg14_atomic_waits/detail/impl/atomic_wait_common.ipp.ipp:173

1.5 atomic_wait_expected_32 timeout error handling is incorrect on Linux

Plan: "Return negative on error."

Actual: wait_on_address32 on Linux always returns -1 on any failure, losing the distinction between ETIMEDOUT and other errors. atomic_wait_expected_32 then returns -1 with errno = 1 (because -ret2 = -(-1) = 1) for any error, including timeout. The code compares ret2 != ETIME && ret2 != ETIMEDOUT, but ret2 is always -1, so the check never matches the positive ETIMEDOUT value.

Files: include/wg14_atomic_waits/detail/impl/atomic_wait_linux.c.ipp:50, include/wg14_atomic_waits/detail/impl/atomic_wait_common.ipp.ipp:653-659


2. Bugs

2.1 Hash table proxy state machine is permanently stuck after first notify

Severity: Critical

Description: The proxy waiter uses proxy->atomic as a binary semaphore:

  • Waiter parks on proxy->atomic == 0
  • Notifier sets proxy->atomic = 1 and calls wake

After any notify, proxy->atomic is permanently 1. It is never reset to 0 for an existing proxy. The proxy is only initialized to 0 when first created in hash_table_find_or_create.

When a subsequent waiter arrives:

  1. It acquires the hash table lock and finds the existing proxy (use_count incremented).
  2. It unlocks and calls the platform wait primitive on proxy->atomic with expected value 0.
  3. Since proxy->atomic == 1 (not 0), the wait returns immediately (EAGAIN on Linux/futex, non-match on Windows/macOS).
  4. The waiter reloads *object. If it still equals expected (which it will if no one notified again), the loop repeats.
  5. Step 3 repeats forever: proxy->atomic is still 1, so the wait returns immediately again.

Result: Infinite busy-loop on Windows/macOS; atomic_wait_generic returns -1 on Linux.

Root cause: There is no mechanism to reset proxy->atomic back to 0 after all waiters have consumed a notify.

Files: include/wg14_atomic_waits/detail/impl/atomic_wait_common.ipp.ipp:38-45, include/wg14_atomic_waits/detail/impl/atomic_wait_common.ipp.ipp:220-222

2.2 atomic_wait_generic returns -1 on proxy spurious wake instead of looping

Severity: High

Description: In atomic_wait_generic, after calling the proxy wait primitive:

const int ret2 = WG14_ATOMIC_WAITS_HASH_TABLE_ITEM_PROXY_TYPE_WAIT(item, ...);
if(ret2 < 0)
{
  return ret2;
}

If the proxy wait returns a negative value (e.g., -EAGAIN because proxy->atomic != 0), the function returns -1 immediately. This is incorrect: a spurious wake or a mismatched proxy state should trigger the re-compare-and-re-park loop, not an error return.

Files: include/wg14_atomic_waits/detail/impl/atomic_wait_common.ipp.ipp:411-418

2.3 atomic_wait_expected_32 returns -1 on proxy error instead of looping

Severity: High

Description: On non-native backends (or when using the hash table fallback), atomic_wait_expected_32 delegates to atomic_wait_generic. If atomic_wait_generic returns -1 due to the proxy state bug (section 2.1), atomic_wait_expected_32 propagates the -1 to the caller. Per the proposal, the only negative return should be on actual system errors, not on spurious proxy state mismatches.

Files: include/wg14_atomic_waits/detail/impl/atomic_wait_common.ipp.ipp:663-664

2.4 wait_on_address32 (Linux) discards errno, preventing timeout detection

Severity: Medium

Description: The Linux wait_on_address32 wrapper:

int ret = (int) syscall(SYS_futex, ...);
if(ret == 0 || errno == EAGAIN || errno == EINTR)
{
    errno = save_errno;
    return 0;
}
errno = save_errno;
return -1;

It always returns -1 on failure, regardless of whether the failure was ETIMEDOUT, EINTR, or something else. Callers cannot distinguish timeout from other errors.

Files: include/wg14_atomic_waits/detail/impl/atomic_wait_linux.c.ipp:43-51

2.5 atomic_notify_generic creates a proxy even when no waiters exist

Severity: Low

Description: atomic_notify_generic calls hash_table_find_or_create(table, object, false). If no proxy exists, this creates one with use_count = 0, stores 1 to proxy->atomic, and calls wake. The proxy is left in the hash table with use_count = 0 and proxy->atomic = 1. When the first waiter later arrives, it inherits the stuck proxy->atomic = 1 state, triggering the infinite loop described in section 2.1.

Files: include/wg14_atomic_waits/detail/impl/atomic_wait_common.ipp.ipp:443-446


3. Race Conditions (Lost Wakes)

3.1 Lost-wake race in hash table proxy between notify and waiter park

Severity: Critical

Timeline:

  1. Thread A calls atomic_wait, loads *object == expected, enters the hash table path.
  2. Thread A is preempted before calling hash_table_find_or_create.
  3. Thread B calls atomic_notify_one (or atomic_notify_all). It acquires the hash table lock, looks up the proxy. If none exists, one is created with use_count = 0. It stores 1 to proxy->atomic and calls wake_by_address32. No one is waiting, so the wake has no effect.
  4. Thread B releases the lock.
  5. Thread A resumes, acquires the lock, finds the proxy (or creates one), increments use_count to 1, releases the lock.
  6. Thread A calls the platform wait on proxy->atomic with expected 0.
  7. Since proxy->atomic == 1 (set by Thread B in step 3), the wait returns immediately.
  8. Thread A reloads *object. If still == expected (which it is, because Thread B may not have changed the value), Thread A loops.
  9. Thread A calls wait again. proxy->atomic is still 1. Wait returns immediately again.
  10. Result: Thread A loops forever (Windows/macOS) or returns -1 error (Linux).

Root cause: The notify path sets proxy->atomic = 1 as a "wake token", but there is no handshake to ensure the waiter has observed the token before it is reused. The proxy has no generation counter or sequence number.

Files: include/wg14_atomic_waits/detail/impl/atomic_wait_common.ipp.ipp:220-222, include/wg14_atomic_waits/detail/impl/atomic_wait_common.ipp.ipp:41-45

3.2 Lost-wake race in native paths between load-check and park

Severity: High

Timeline (native Linux/macOS/FreeBSD/Windows):

  1. Thread A calls atomic_wait, loads *object == expected.
  2. Thread A is preempted.
  3. Thread B calls atomic_notify_one without changing *object (e.g., notify_all in a scenario where the value is already the desired state).
  4. Thread B calls the platform wake primitive (FUTEX_WAKE, __ulock_wake, etc.) on object. No thread is currently parked, so the wake has no effect.
  5. Thread A resumes and calls the platform wait primitive (FUTEX_WAIT, __ulock_wait, etc.) on object with expected value.
  6. The kernel parks Thread A because *object == expected.
  7. No further notify is forthcoming. Thread A blocks forever.

Why the while-loop doesn't help: The while-loop (while(load == expected) { wait(); }) only helps if the wait returns spuriously or because the value changed. In this race, the wait blocks because the value is still expected, and no one will wake Thread A.

This is a well-known limitation of futex-style wait/notify when the notify does not change the value. The C++20 atomic_wait / atomic_notify design assumes notify is paired with a store that changes the value, but the proposal does not require this.

Files: include/wg14_atomic_waits/detail/impl/atomic_wait_linux.c.ipp:43-44, include/wg14_atomic_waits/detail/impl/atomic_wait_macos.c.ipp:64-66, include/wg14_atomic_waits/detail/impl/atomic_wait_common.ipp.ipp:494-499


4. Missing Tests

4.1 No tests for hash table path

The existing tests use atomic_int, which is 4 bytes. On Linux, macOS, FreeBSD, and Windows, 4-byte operations bypass the hash table. No tests exercise 1-byte, 2-byte, or 8-byte (Linux) operations that fall back to the hash table.

Impact: The critical bugs in sections 2.1, 2.2, and 3.1 are not caught by CI.

4.2 No tests for notify-without-store race

The tests always store a new value before calling notify. They do not test the scenario where notify is called without changing the value, which triggers the lost-wake race in section 3.2.

4.3 No tests for multiple waiters on the same object

atomic_notify_test.c creates multiple waiters but uses atomic_notify_all to wake them all. It does not test atomic_notify_one with multiple waiters, or the case where max_threads_to_wake limits the number of woken threads.

4.4 No tests for spurious wakes in hash table path

The tests do not inject spurious wakes into the hash table proxy to verify that the re-compare-and-re-park loop works correctly.


5. Other Issues

5.1 atomic_wait_expected_32 timeout conversion uses ceiling, but wait returns -1 on all errors

The timeout computation in atomic_wait_expected_32 correctly uses ceiling conversion (tv_nsec + 999999) / 1000000). However, because wait_on_address32 returns -1 for all errors (section 2.4), the caller cannot distinguish a real timeout from a spurious wake or other kernel error. The outer loop re-checks the deadline before each park, which mitigates this somewhat, but the return value semantics are still wrong.

Files: include/wg14_atomic_waits/detail/impl/atomic_wait_windows.c.ipp:52-53, include/wg14_atomic_waits/detail/impl/atomic_wait_macos.c.ipp:58

5.2 atomic_wait_common.ipp.ipp default proxy uses uint_least32_t but atomic prefix is missing

The default proxy definition:

#define WG14_ATOMIC_WAITS_HASH_TABLE_ITEM_PROXY_TYPE                           \
  WG14_ATOMIC_WAITS_ATOMIC_PREFIX atomic_uint_least32_t

The ATOMIC_PREFIX is applied, which is correct for C (atomic_uint_least32_t) but in C++ it becomes std::atomic_uint_least32_t. This is correct because the header includes <atomic> in C++ mode.

5.3 hash_table_remove_item doesn't reset proxy->atomic before free

Not a bug per se (the proxy is freed), but if the removal is ever skipped due to a non-zero use_count, the proxy remains with proxy->atomic = 1.

5.4 hash_table_grow doesn't reset old bucket keys to NULL after rehashing

In hash_table_grow, after rehashing, the old buckets are freed with free(table->buckets). The new buckets are freshly calloc'd, so they are zeroed. This is correct.

5.5 atomic_notify_32 fast-path for max_threads_to_wake == 0 returns 0 without CAS

The plan explicitly allows this:

"Use a fast-path check before acquiring the hash table lock: if max_threads_to_wake == 0, skip the hash table lookup and wake_waiters call entirely and return 0 immediately (no CAS, no notify, no side effects)."

Files: include/wg14_atomic_waits/detail/impl/atomic_wait_common.ipp.ipp:674-677

5.6 atomic_wait_generic sets order = success unconditionally after wait returns

If the wait returns because of an error (e.g., -EAGAIN from the proxy), order is still set to success. This means the subsequent load uses success ordering even though no notify occurred. The proposal requires failure ordering for accesses before any notify.

Files: include/wg14_atomic_waits/detail/impl/atomic_wait_common.ipp.ipp:419

5.7 atomic_wait_expected_32 return value on timeout is 0, but the proposal says "zero if no suspension occurred, or if duration timeout occurs"

This is correctly implemented: the function returns 0 on timeout and *expected is updated to the timed-out value.

Files: include/wg14_atomic_waits/detail/impl/atomic_wait_common.ipp.ipp:631-636


6. Summary of Critical Issues

Issue Severity Location
Hash table proxy stuck at atomic=1 after notify Critical atomic_wait_common.ipp.ipp:220-222
Lost-wake race: notify before park in hash table path Critical atomic_wait_common.ipp.ipp:41-45
atomic_wait_generic returns -1 instead of looping on proxy error High atomic_wait_common.ipp.ipp:411-418
Lost-wake race: notify without store in native path High atomic_wait_linux.c.ipp:43-44
Windows notify_all uses WakeByAddressSingle Medium atomic_wait_windows.c.ipp:100
Windows 1/2-byte ops use hash table instead of bypass Medium atomic_wait_common.ipp.ipp:524-528
wait_on_address32 discards errno Medium atomic_wait_linux.c.ipp:50-51
macOS notify_all ignores max_threads_to_wake limit Medium atomic_wait_macos.c.ipp:99-106
No tests for hash table path High test/atomic_wait_test.c
No tests for notify-without-store High test/atomic_notify_test.c

The most severe issue is the hash table proxy state machine (section 2.1 / race 3.1). The proxy is a single-slot binary semaphore that is set to "notified" (1) by the notifier but never reset to "waiting" (0) after consumption. This makes the hash table path fundamentally broken for any object that receives more than one notify, or for any waiter that arrives after a notify has already been issued for that object.

Deepseek v4 Flash (0731) — Implementation Review

Date: 2026-08-01 Scope: Reviewed the wg14_atomic_waits implementation against docs/proposal.md (the authoritative spec) and plans/wg14_atomic_waits-reference-library.md. Method: Static inspection of every include/ backend, src/ wrapper, the public header, config.h, the tests, the build files, and CI. No debugger was run and no test program was executed; all conclusions come from reading the code.


1. Executive summary

The implementation is structurally faithful to the proposal (all six public APIs exist, width dispatch for 1/2/4/8-byte atomics, native wait/notify backends with a hash-table fallback), but it contains one primary correctness race that produces lost wakes (Section 2), a broken pthreads backend (Section 3), a timeout smuggled in as an error on Linux (Section 4), a Windows lost-wake for multi-thread wakeups (Section 5), plus a number of smaller deviations from the plan/proposal.

The single most important finding: the per-object "proxy" used by the hash-table path sets a 0/1 notification flag that is never cleared while the wait-queue node is alive. This is the lost-wake/livelock race the task asked to identify.


2. PRIMARY RACE: the proxy notification flag is never reset → lost wakes / busy-spin

Location

include/wg14_atomic_waits/detail/impl/atomic_wait_common.ipp.ipp:

  • ..._WAIT macro (lines 37–38):
    wait_on_address32(&(x)->atomic, 0, (timeout))   /* wait while proxy->atomic == 0 */
    
  • ..._WAKE macro (lines 39–45):
    atomic_store_explicit(&(x)->atomic, 1, release),  /* mono-directional: 0 -> 1 only */
    wake_by_address32(&(x)->atomic, max_threads_to_wake)
    
  • atomic_wait_generic() (lines 327–435) — the shared "park by proxy" loop.
  • atomic_notify_generic() (lines 437–456) — the shared "set flag + wake" path.

The bug

A waiter parks by calling WAIT(item,...) which blocks while item->atomic == 0 (i.e. FUTEX_WAIT(&item->atomic, 0)). A notifier sets item->atomic = 1 and wakes. There is no code anywhere that ever writes item->atomic back to 0 while the wait-queue node is alive. The only place it is reset is at node creation inside hash_table_find_or_create() (lines 220–222), which happens only when a brand-new proxy_waiter_t is allocated. A node is freed only when use_count drops to zero.

Consequence — two interleaving outcomes

(a) Re-park after a wake never sleeps (livelock). Once any notify has fired on a node, item->atomic is stuck at 1 for as long as the node lives. Any waiter that is woken and must re-park — the proposal explicitly requires re-park on spurious wake, and the code implements it as the top of the loop — calls FUTEX_WAIT(&item->atomic, 0) while the value is already 1. The kernel compares 1 != 0 and returns EAGAIN immediately. Every subsequent iteration of the loop returns immediately, so the waiter never sleeps again; it degenerates into a tight 100%-CPU busy-spin for the whole remaining lifetime of that node.

(b) Wakeups are lost because there is no sleeping thread. Because (a) means waiters stop sleeping, a later genuine producer store + notify_* sets the flag (already 1) and issues FUTEX_WAKE, which has nothing asleep to wake. The notify is therefore effectively lost for the purpose of the sleep/wake contract; correctness then depends entirely on the busy-spin poll observing the value change, which is not the semantics the proposal defines and not what a correct reference implementation should do.

Why the analogous futex idiom would be safe but this one is not

The correct pattern guards the "am I allowed to sleep" decision on the same state that the notifier flips, and the notifier re-arms the state before waiting:

  • waiter: s = counter.load(); if (value == expected) futex_wait(&counter, s);
  • notifier: counter++; futex_wake(...) Here counter is a strictly increasing generation so the waiter can always detect a change that happened between its load and its sleep. The implementation instead uses a single 0/1 flag that is never re-armed, so the invariant "atomic == 0 ⇔ a notify is pending/expected" is destroyed after the first notify.

Which configurations suffer

This path is the fallback for every backend whenever the operand cannot be handled directly by the kernel primitive, i.e. exactly the cases the proposal/plan force through the hash table:

  • Linux: 1-, 2- and 8-byte operands (HAVE_WAIT_ON_ADDRESS_* is 32-bit only).
  • macOS / FreeBSD: sub-native widths (1/2-byte).
  • pthreads backend: every operand (there is no kernel per-address waiter).

The 4-byte Linux/macOS/FreeBSD/Windows fast paths and the atomic_wait_expected native-width path bypass the proxy and are not affected by this flag, but the 8-byte-on-Linux case — a perfectly legal and likely test target — is affected.

Replace the 0/1 flag with a monotonically increasing sequence number that the waiter reads before parking and passes as the futex compare value, and that the notifier increments before waking. Reset-on-rearm must happen on the waiter side before the sleep decision, under the same lock used to re-check the object value (or rely on the kernel re-check for the object value itself as the futex fast path already does).


3. pthreads backend is fundamentally broken (hangs / lost wake)

include/wg14_atomic_waits/detail/impl/atomic_wait_pthreads.c.ipp:

  • ..._WAIT (lines 45–46) is pthread_cond_wait(&(x)->atomic, pthreads_mutex()).
  • pthreads_mutex() (lines 61–73) returns a _Thread_local mutex, i.e. a different mutex object per thread.

Problems:

  1. pthread_cond_wait requires the passed mutex to be held by the calling thread. In atomic_wait_generic the waiter has released the hash-table lock (line 387) and then enters pthread_cond_wait with a mutex that is never locked. This is undefined behavior; on glibc it typically fails immediately (EPERM) so the wait “succeeds” without ever blocking — again a busy-loop — and there is no guarantee the node is protected.
  2. The broadcast hand-off is not protected by the mutex the waiter sleeps on. A notifier holds the global hash-table lock and calls pthread_cond_signal (via the ..._WAKE macro, lines 47–53). The classic lost wake occurs when the notifier signals between the waiter’s re-check (value still equal to expected, line 380) and its pthread_cond_wait: the signal is dropped and the waiter blocks forever. With a futex, the kernel’s value re-check/EAGAIN saves this; with pthread_cond_t there is no such guard and there is no predicate/flag protecting the check, so the wait is a genuine, permanent lost wake (a hang).
  3. Even the INIT/DESTROY macros treat pthread_cond_t through the generic proxy_waiter_t.atomic slot, but the shared atomic_wait_generic still performs flag-style logic (setting use_count, etc.) that is meaningless for a condvar.

Because CI runs ALWAYS_USE_PTHREADS_BACKEND=ON on Linux and macOS, this path is exercised, but its crashes/hangs are exactly the class of lost-wake bug being reported.


4. atomic_wait_expected mis-reports a timeout as an error on Linux

include/wg14_atomic_waits/detail/impl/atomic_wait_common.ipp.ipp, atomic_wait_expected_32() (lines 595–666), plus the Linux wait_on_address32() (atomic_wait_linux.c.ipp lines 37–52).

  • The Linux wait_on_address32 returns 0 on success/EAGAIN/EINTR and -1 on any other error, not -errno. A genuine time-out of FUTEX_WAIT therefore comes back as -1 (with errno == ETIMEDOUT).
  • The caller’s error branch (lines 651–659):
    if(ret2 < 0)
    {
      if(duration != NULL && ret2 != ETIME && ret2 != ETIMEDOUT) { errno = -ret2; return -1; }
    }
    
    ret2 is -1, which is never equal to the positive ETIME/ETIMEDOUT constants, so the condition is always true for any ret2 < 0 when a duration was supplied. A clean time-out returns -1 (error), not 0 (time-out) as the proposal requires:

    Returns: … returns zero … or duration timeout occurs.

This is timing-dependent — if the pre-wait clock_gettime check (lines 631–637) happens to notice expiry first it returns 0 cleanly — so the 1 ns test in atomic_wait_expected_test.c is flaky, but the underlying error path is wrong.


5. Windows wake_by_address* only ever wakes a single thread → lost wake

include/wg14_atomic_waits/detail/impl/atomic_wait_windows.c.ipp, wake_by_address32/wake_by_address64 (lines 92–124):

if(WakeByAddressSingle((PVOID)(uintptr_t) object)) return (max_threads_to_wake == 1) ? 1 : 1;
return 0;
  • The max_threads_to_wake parameter is ignored; both atomic_notify_all and atomic_notify(..., max_threads_to_wake=N>1, ...) call this and wake exactly one thread via WakeByAddressSingle. The correct routine for max != 1 is WakeByAddressAll. Every other waiting thread is left parked → lost wake.
  • The (max_threads_to_wake == 1) ? 1 : 1 ternary is dead code (both branches are 1).

This makes the Windows backend incorrect for atomic_notify_all and for atomic_notify with max_threads_to_wake > 1, which the plan marks as mandatory behaviour.


6. macOS timeout conversion deviates from the plan

atomic_wait_macos.c.ipp, wait_on_address32/64 (lines 55–68, 77–90):

  • The plan (Step 11) requires: *durationnanoseconds, cap each ulock_wait call at UINT32_MAX (~4.29 s) and loop for longer durations.
  • The implementation instead converts once to microseconds (tv_sec * 1000000U + tv_nsec / 1000U) and passes it in a single call with no cap and no loop. For any duration ≥ ~4295 s the microsecond value overflows uint32_t, and durations beyond ~4.29 s are not split across multiple calls, so the accumulated wait can be far shorter than *duration — violating the proposal’s “total accumulated time … shall be at least *duration”.
  • The code also declares private extern __ulock_wait/__ulock_wake instead of including <bsd/sys/ulock.h> as the plan directs; functional risk if SDK/version behaviour differs.

7. FreeBSD 8-byte UMTX_OP_WAIT argument-order inconsistency

atomic_wait_freebsd.c.ipp:

  • 4-byte: _umtx_op(object, UMTX_OP_WAIT_UINT, expected, (long)&umtx_time) (lines 57–58) — passes the expected value in the value slot.
  • 8-byte: _umtx_op(object, UMTX_OP_WAIT, (long)&umtx_time, (long)expected) (lines 88–89) — passes the timeout pointer in the value slot and the expected value in the address slot, i.e. the two are swapped relative to the 4-byte call.

The two calls are internally inconsistent, so at least one passes the operands in the wrong order; UMTX_OP_WAIT (8-byte) is almost certainly wrong and will compare against garbage / misbehave.


8. Return-value deviations from the plan

  • atomic_notify_32 (lines 668–691) returns 1 + ret on a successful CAS, where ret is the number actually woken (0 on the proxy path when no node exists, 0 on the futex path when nothing is parked). So CAS-success-with-no-waiters returns 1 (positive), whereas the plan explicitly states:

    atomic_notify Returns 0 if the CAS fails or no waiters are parked.

    (The proposal’s “possibly one plus the number woken” makes 1 defensible, so this is a plan deviation, not a proposal violation — noted for completeness.)

  • atomic_wait_generic sets ret = 1 (line 388) before the first actual park call. The proposal ties “positive” to “suspended at least once”; the intent-to-park flag is acceptable but slightly loose.

9. Width-dispatch macros silently no-op on unsupported widths

atomic_wait.h, _WG14_ATOMIC_WAITS_IMPL_atomic_wait* / _notify* (lines 107–201):

  • Each macro is if (sizeof==1) … else if (==2) … else if (==4) … else if (==8) … with no final else (and do{}while(0)). A _Atomic type of any other width (e.g. 16 bytes, or a long double, or a 0-width type) compiles to a silent no-op rather than a compile-time error, which can mask misuse. The plan states widths 1/2/4/8 are supported; a diagnostic would be safer.

10. Header-only / ODR notes

  • The atomic_wait_*_N / atomic_notify*_N definitions in the .ipp files are not themselves marked WG14_ATOMIC_WAITS_INLINE/static (only their prior declarations in atomic_wait.h carry WG14_ATOMIC_WAITS_EXTERN, which is inline only when WG14_ATOMIC_WAITS_ENABLE_HEADER_ONLY). In the multi-TU header-only build (header_only_test*.cpp) this relies on the C inline extern-inline model. It is workable on GCC/Clang but fragile; the plan required every backend helper to be explicitly static inline to guarantee ODR safety.
  • hash_table() uses WG14_ATOMIC_WAITS_IGNORE_MULTIPLE_DEFINITIONS (weak), so the singleton table is fine across TUs; this part is OK.

11. Smaller issues

  • errno not preserved on success/failure paths of atomic_wait_expected_32: it writes errno = ETIMEDOUT (line 636) and errno = -ret2 (line 656), clobbering the caller’s errno even though errno preservation was a stated plan requirement (Steps 8–12). The futex wait_on_address32 does restore errno, but the caller overwrites it afterwards.
  • EINTR treated as success (return 0) in wait_on_address32 (Linux line 45) rather than retrying; a signal during a timed wait silently shortens it.
  • atomic_notify_generic drops a notify when no node is registered (find_or_create with increment_use_count=false, lines 444–445). For the futex backends this is masked by the waiter’s under-lock re-check; for the pthreads backend it is not masked (see Section 3), which is where the drop becomes a hard lost wake.
  • Hash-table growth/deletion re-probes correctly (cluster-shift deletion, lines 271– 280), but hash_func truncates a 64-bit pointer to 32-bit before mixing (lines 108–113); not a correctness bug, just avoidable collisions.

12. Structure deviations from the plan (non-bug)

  • The plan (Step 7) specified that atomic_wait/atomic_wait_explicit be thin wrappers delegating to atomic_wait_expected, and notify_one/notify_all delegate to a wake_waiters helper. The implementation instead uses per-width atomic_wait_1/2/4/8 + direct kernel calls for the native widths, with the hash-table path only for the fallback. This is a reasonable and arguably more correct choice given that atomic_wait_expected is native-width-only, but it is a deviation from the plan’s described structure.
  • benchmark_atomic_wait_test.c does not actually run a benchmark (returns 0 after declaring an unused consumer), so CI excludes it correctly but it validates nothing.

13. Conclusion

The implementation gets the public API shape, the width dispatch, and the native-width futex path right, but the hash-table/fallback proxy logic has a one-way notification flag that is never re-armed, which is the race responsible for lost wakes: waiters that re-park on a node that has already been notified stop sleeping and busy-spin, and subsequent notifications have no sleeping thread to wake. On the pthreads backend the same logic is additionally broken by a thread-local mutex and a missing predicate under pthread_cond_wait, producing hard lost-wake hangs. These are compounded by a Linux timeout→error bug, a Windows single-wake bug for notify_all/max>1, a macOS timeout conversion that violates the minimum-duration guarantee, and a likely FreeBSD 8-byte argument swap. All were identified by code inspection only.

Scroll down to see more ...

Differences:

  1. Only Step 3.7 noticed that the Windows backend doesn’t bypass for all atomic sizes as it should according to the implementation plan.
  2. Step 3.7 false positives ‘atomic_notify_generic creates a proxy even when no waiters exist’.
  3. Step 3.7 side mentions the lack of a generational count: Deepseek v4 gets properly into the issue and describes it in full and does it well at that.
  4. Deepseek v4 has a much better handle on what’s broken in the pthreads backend and why.
  5. Only Deepseek v4 spots the inverted arguments in the FreeBSD backend.

Also, in general, the Deepseek v4 report reads more coherently and gets into the big issues quickly and tersely, whereas the Step 3.7 report is bitty and kinda all over the place.

Neither did a good job of identifying where the implementation plan or the implementation deviate from the proposal. The implementation plan specifically states at its top:

docs/proposal.md is the authoritative spec. Behavior, return values, and memory-order semantics must match it exactly.

After this I stopped using Step 3.7 and exclusively used Deepseek v4. Perhaps the latter was overwhelmed by all the defects and deviations from specification with this first analysis which is why it didn’t perform well – all I can say is that later on, perhaps as the implementation’s major bugs got fixed which made logic analysis easier, Deepseek v4 Flash began to seriously impress me with its analysis capabilities. One still has to go through multiple rounds of something like:

Exhaustively compare the implementation to the proposal, looking for all cases of deviation, bug, issue, concerns and corner case where the implementation does not match the proposal, or the proposal may not be implemented in full. Be very detailed, thorough and methodical in your approach - go that extra mile each and every time. Analyse in depth all implementation backends and all combinations of build configuration options, also analyse in depth all error handling and control flow paths not normally taken. Update plans/combined-analysis with your findings.

… and then you rinse and repeat iterations of that, fixing one by one all the things it finds, and doing so until it finds nothing important in your opinion. So in that sense it’s like Step 3.7, but where it massively improves is on one shot fixes for CI failures: you simply tell it which CI failed and copy and paste the failure text. It then had a 100% success rate at finding and fixing the CI failure even for platforms it could not debug locally – it did so simply by inspection and analysis, including inspecting online the kernel sources for Linux, FreeBSD or ReactOS (to get an insight into Windows).

In fact, there was an especially impressive bit where it found that Apple Clang 17.0 only would produce invalid binaries if symbol visibility was set to hidden and a specific tail call optimisation caused a function to be inlined into main(). It went off decompiled the Apple Clang binary, compared it against the LLVM clang source code, found the exact bit of problematic reverse compiled source code, wondered to itself it it ought to patch the Apple Clang binary, and eventually decided to instead hack around the problem for this specific instance and it even added an informative explanatory comment to say why its hack was there.

Now, I’d read of Claude doing stuff like that. I’d seen Step 3.7 analyse the assembler in binaries to figure out why they weren’t performing as expected. But to actually localise a bug in a third party precompiled binary via reverse engineering? That was new to me. No doubt it did take rather a long time to do all that – due to Deepseek v4’s immense popularity right now, it has not been running quickly, as little as 30 toks/sec. But I could leave it chug away on its own safely I found, whereas Step 3.7 had a nasty habit of occasionally wrecking your git repo or going off and installing huge bits of software it didn’t really need via brew.

After the WG14 atomic waits reference library was finished, I had spent US$5.02 on 465 million tokens. That is US$0.0108 per million tokens. Yes that is an awful lot of tokens – in fact, I have consumed 913 million tokens ever on OpenRouter, so this one project consumed half my lifetime total – but the quality of implementation created is very high in my opinion. I would estimate it would have taken me over one hundred hours to create a similar quality implementation by hand – instead this cost me less than ten hours in total, which was almost entirely spent reviewing its work and giving direction on what to do next. Five dollars for ninety hours of my life back to do more interesting work is a bargain.

Test 3: Subjective experience of using each LLM to get work done

I think it’s fair to say I’ve been repeatedly wowed by Deepseek v4 Flash 0731’s capabilities. I HAVE found that you should not let it take architecture direction decisions: always ask it to present a menu of implementation options, and you’ll find half the time its recommended implementation is the wrong one. So that part sucks. But when you choose on its behalf the right implementation option, 98% of the time it does a great job: it matches the style and form of the existing codebase, it avoids writing copy and paste code and instead hoists common routines into reasonable locations in reasonable common header files, and the code quality written is well above most of the programmers I’ve ever worked with, with only very occasional slip ups. I really like the much improved one-shot fix capability, especially for platforms and architectures I can’t run on my system where I have the LLM agentic harness running. That’s been a HUGE timesaver: no more having to boot up Windows VMs etc to diagnose some random failure on CI.

I very much like its performance analysis. I asked it to make this codebase go faster. It spent some minutes pondering and reading code, and it told me we ought to use triangular probing instead of quadratic probing in the open addressed hash table as the buckets are a power of two, so the triangular probing would ensure better scattering of entries avoiding collisions. It one-shotted the new implementation, then benchmarked the difference, then twiddled a few unrelated items by parsing through the optimised disassembly as it knew my main ask was for improved performance. It then spat out hard benchmarks: 470 nanoseconds reduced to 50 nanoseconds. Impressive. It also generated comprehensive tests for scalability under load and that bucket growth did work perfectly under heavy multithreaded load. Even more impressive.

I asked it to add the Fil-C toolchain to the CI (this is a guaranteed memory safe C/C++ toolchain). It went off and found the documentation on the web, followed the instructions, set up the appropriate Github CI actions, adjusted the codebase where necessary as the Fil-C libc is musl rather than glibc, then to test it it installed a Linux VM as this is a Macbook, installed the correct AArch64 edition of Fil-C rather than the x64 one the CI uses, and ran the test suite via the local Linux VM under Fil-C. Worked first time in a single shot too – it didn’t make a single mistake. Had I done that by hand, I definitely would have made a mistake at least once – I know from past experience that setting up the Fil-C toolchain is finicky.

I originally had asked Step 3.7 to create this new reference library using an existing hand written reference library as its template. It didn’t do too well at that, and with hindsight I wish I’d have wiped what it did and started from scratch with Deepseek v4. Deepseek v4 Flash makes far fewer mistakes within the test harness (Kilo Code) and doesn’t need to self correct anything like as much. It also gets tool calling in the Kilo code harness right almost all of the time, unlike Step 3.7. I suspect it would have done a better job at mimicking the template into this new library, but I guess I’ll not find out until the next time I write a new reference library for WG14.

Deepseek v4 Flash is much more prone to proactively fix bugs and issues without being explicitly told it can first. It seems to ask for forgiveness rather than permission. So you’ll need to be careful to always git commit before asking it a question, otherwise it may decide your question demands code changes. At least it doesn’t like to git reset --hard as Step 3.7 Flash was keen on doing when it got the codebase into a confused state, rather it properly uses git stash so any working tree changes can be recovered.

The 1M max context of Deepseek v4 Flash makes a BIG difference! I was used to beginning to sweat as the 260k context limit approached, trying to get it to write out todo lists into Markdown files for the next session clear as I always found context compaction just didn’t work well with Step 3.7. With Deepseek, I can just relax and let it trundle on – you are still wise to start a new session from time to time as the long contexts slow its execution down, but now you can take your time about it, and more importantly, if it goes off on a long extended think or ponder or diagnosis of something you can just ignore it because it won’t suddenly run out of context. This is the first occasion I can go do some other task while it runs in the background and I don’t need to stress constantly checking its progress. Very nice!

And finally, I really like how much cheaper it is. On Openrouter comparing usage and spend from before to after, I issued twice the requests and spent half as much money. I didn’t have to babysit this model as much as before. All in all, this is my new favourite Pareto cost-benefit optimum LLM choice. Well, at least for the next three months, if the pattern so far this year continues to hold true!

Which brings me onto …

Where LLMs and AI probably are going next

We now have enough history of LLM evolution to be able to predict with reasonable reliability where things will go next. As I mentioned above, I’ve chosen a new LLM coding assistant every three months on average this year. Shall I continue to do so?

I had Deepseek v4 Flash go off and scrape the AA Intelligence Index for a spread of LLMs over the past two years off https://artificialanalysis.ai/, then plot those using a contour map:

This has contour bands for ventiles in a LLMs AA Intelligence Index score and it shows that:

  1. For a <= 10 billion parameter model which is feasible for me to run locally given my ancient hardware, we broke into the tens around January this year, and we would expect to break into the twenties any time around now, followed by the thirties around January 2026, and the forties around Summer 2027. So, by Summer 2027, a 10 billion parameter model will score as well as Deepseek v4 Flash does. And it’ll run well even on this old Apple M3 based laptop.
  2. For a <= 300 billion parameter model which is likely to run well on near future Apple Macbook Pros, we broke into the twenties around last November, into the thirties last March, into the forties last week, and we would expect to break into the fifties before the end of 2026, then into the sixties by March 2027. Reminder: Claude Fable 5 scores sixty-one. So, before the end of Spring 2027, that six trillion parameter model will score similarly to a 300 billion parameter model!

Let’s graph time directly against AA Intelligence Index score:

What strikes you about this graph is firstly by how much the open weights models are catching up with the closed weights ones – I would be surprised therefore if the Chinese government continue to release their frontier models with downloadable weights in the near future. Secondly, there is a clear structural break between 100-200b and 200-500b models – there is wide space between their trend lines. Indeed, right now 200-500b is outperforming 500-1T, which is surprising. Less surprising is the 1T+ category which has a trend line matching that of the closed weight models.

The current most intelligent LLM anywhere by this index as of August 2026 is Claude Opus 5 with a score of sixty-one, followed by the current most intelligent free to download LLM which is Kimi K3 with a score of fifty-seven (Kimi K3 is a cool 1.4 Tb of download, and generally you need about as much VRAM as the download size to run it, so that would be very expensive to run locally given current RAM prices). Deepseek v4 Flash 0731 is 167 Gb to download, so it would run well on a machine with 256 Gb of VRAM, and it gets an intelligence score of 49.9. Of course, a single score is an average, and some models are strong and weak on specific domains compared to others. https://artificialanalysis.ai/ lets you compose comparisons of LLMs, so I chose these six as representative for this discussion:

Scroll down to see more ...

As you can see, the AA Intelligence Index score is made up of lots of separate indices, each of which is then weighted into an average overall score. On some specific domains e.g. GPQA Diamond or r3-Banking, already it’s a wash between recent models. On some others, there is a clear pattern of the gap rapidly narrowing, however there are always going to be things at which a five trillion parameter model will beat the pants off a 500 billion parameter model: tool use and logic aren’t those, but specialist knowledge and reasoning will be.

In other words, yes while the AA Intelligence Index score will improve over time for smaller models, that will be in those parts of the index which aren’t specialist knowledge and reasoning. Small models simply can’t store as much knowledge as larger ones, so I expect my estimations above to be rather optimistic i.e. the score improvements will be less for the smaller parameter models than one would currently predict by extrapolation from recent past. Especially because the AA Intelligence Index score is out of one hundred, so as models max out all parts except the specialist knowledge and reasoning, they will end up running into an upper bound where only more parameters can improve some of their domain specific scores.

ALSO all this prediction is contingent on the AI investment bubble continuing to inflate. One gets cleverer small models by investing more compute into training fewer parameters. For that, you need more and cheaper compute, and for that you need to keep investing those billions. This year they think ~US$900 billion has been invested in AI, and for next year we are on track for US$1.4 trillion dollars in 2027. Even the great wealth and income of the tech multinationals will struggle to fund so much debt – even now, total free cash flow for several of them is below their debt servicing costs for the debt they’ve taken out. But that’s another diary entry. In any case, it is hard to believe that the AI investment bubble won’t pop soon, and then we’ll have whatever compute has been built out by then and it’ll only grow linearly rather than exponentially after that, much like with the late 1990s telecommunications infrastructure investment bubble.

Linear compute growth does still enable model improvements, and unsurprisingly I’d expect them to stop improving exponentially and start improving linearly instead. As with the end of Moore’s law, the slowdown will affect the biggest highest end models first, and the smaller lower end models will see a long run of continuing exponential growth before that eventually also peters out. It’s been the same with CPUs: at the very cheap end, they’ve been continuing to exponentially improve the value per dollar cost for decades after the high end went into linear improvements. I think the same will apply to LLMs: after all, if training cost per dollar goes from exponential to linear improvement, the lessons learned from making the high end a little better should translate into larger improvements at lower ends, same as for CPUs.

I find this prediction of the future FAR more believable than predictions of imminent Technological Singularity which have started doing the rounds again. I covered that in the unpublished book I wrote after St. Andrews: the Singularity is purely the result of an artefact of human perception where we tend to weigh more recent big leaps forward as more important, as they are more important to us personally but aren’t really in the bigger picture of things outside humanity. Elon Musk had an interview with the Economist week before last where he was banging on about the Singularity. I suppose that suits his purposes to market that philosophy aggressively so fewer think about seizing some of his trillion dollars of personal wealth, but I also got the impression from the interview that he actually genuinely believes that a Singularity will happen at some point. I’ll categorically state right now: no Technological Singularity will happen in my lifetime unless some very new technology turns up. Certainly nothing about Large Language Models as presently designed and implemented is capable of generalised artificial intelligence i.e. AGI. Right now we’re in the exponential growth phase because we’re pouring exponential amounts of capital in – cut the constantly increasing capital investment flows and you can say good bye to exponential LLM capability improvements, as I just described above. All that said, the near term improvements to consumer hardware WILL be significant to Economic Total Factor Productivity as the gains from this technological advancement begin to diffuse widely throughout society.

Near future hardware

So that brings me onto the near future consumer hardware to run these things. Recent leaks say that Apple have started to design their next M-series and A-series chipsets to have better than the usual trendline of improvements to compute and memory bandwidth, so your 2028 Apple Macbook Pro should locally run < 500 bn parameter LLMs quite well indeed. Here are the current rumours and leaks in a single table and graph as created from the table by Deepseek v4 (at which it was surprisingly poor at doing interestingly, I really had to poke it hard and repeatedly to generate correct looking SVG, despite it amazing performance at graph building shown above – maybe the HTML table input upset it?):

Model / Year Edition GPU Cores Memory Bandwidth Remarks
Apple M3 2023 Pro 18 cores 154 Gb/sec Unfortunately my personal Macbook is the M3 Pro, the worst for running LLMs of any of the Pro Macbooks 🙁
Max 40 cores 410 Gb/sec
Apple M4 2024 Pro 20 cores 273 Gb/sec Added a memory cache shared between CPUs and GPUs like AMD's Infinity Cache for its GPUs. This greatly improved latency.
Max 40 cores 546 Gb/sec
Apple M5 2026 Pro 20 cores 307 Gb/sec First with hardware matrix multiply and accumulate (= nVidia 'tensor cores'). LLM input parsing is approx 4x faster than M4 as a result. 1024 FP16 FMAs per core per cycle enables 70 FP16 TFLOPs for the Max edition.
Max 40 cores 614 Gb/sec
Apple M6 2027? Pro 32? cores 512? Gb/sec Expected move to LPDDR6 standard memory architecture featuring a wider 24-bit channel layout (shifting to 384-bit Pro / 768-bit Max buses) to achieve a projected 1.67x generational leap in raw memory speeds. There will be no Max nor Ultra edition of the M6, this suggests that the core will be very similar to the M5 and they only upgrade the memory bandwidth.
Max 64? cores 1024? Gb/sec
Apple M7 2028? Pro 48? cores 800? Gb/sec Rumours say the Max variant can be fitted with up to 768 Gb of RAM in your standard Macbook laptop chassis. Obviously so much RAM will be VERY expensive as Apple likes to charge steeply for additional RAM. It would be surprising if TFLOPs don't double due to implementing 2048 FP16 FMAs per core per cycle.
Max 96? cores 1600? Gb/sec

If you extrapolate out the numbers, the Apple M7 Pro should have the same memory bandwidth as a nVidia Volta enterprise AI accelerator from 2017, and the M7 Max should have the same memory bandwidth as a nVidia Ampere AI accelerator from year 2020. Chances are that the Macbook Pro and especially Max will have more VRAM (or equivalent, see below), but in terms of compute with 1024 FP16 FMAs per core per cycle they should pretty much match a Volta and Ampere exactly: the Volta maxed out at 125 FP16 TFLOPs and the Ampere 312 TFLOPs. Both had hardware matrix multiple and accumulate, same as the Apple M-series from the M5 onwards. As mentioned in the table above, it would be surprising if the M7 doesn’t implement at least 2048 FP16 FMAs per core per cycle given that today’s nVidia Rubin chipset can do 16384 FP16 FMAs per core per cycle, and Apple tends to follow closely whatever architecture choices nVidia makes – the M5 chipset’s GPU looks awfully like a nVidia GPU, just less wide. This architectural closeness is why LLM software support tends to be nVidia first, then Apple, then AMD (which is architecturally different), then Intel (which is architecturally different again). And why LLM software support on Apple is first class, whereas although support for AMD has improved enormously, it remains second class.

There are zero rumours about this next bit, so it’s probably wrong, but I would wonder if Apple would fit so much DRAM when flash mounted as Storage Class Memory (SCM) is (i) cheaper and especially (ii) much less drain on battery life. There is zero good reason why LLMs are stored in DRAM other than there isn’t an easily available cheaper substitute, but somebody big like Apple could simply fit NAND flash where the DRAM goes. You might only write that flash with an updated LLM every few months so its endurance won’t matter, and NAND flash if mounted like RAM is nearly as fast as DRAM. As the LLM model weights aren’t mutated in RAM, this could save easily 80% of the RAM demands of a LLM, so you get to run your 200 Gb sized LLM in 40 Gb of DRAM and probably less if you shrink the size of the KV cache which is very doable if you have Ampere levels of compute on tap.

Obviously that’s pure speculation, but I do know that DRAM is hard on battery life as it must be continually refreshed. Storage class memory would be easy to fit for somebody big like Apple and it would fix the battery life impact problem. I guess we’ll find out in 2028. In any case, you would expect parsing of around four thousand tokens per second, and generation of a hundred tokens per second on the M7 Pro – and double that for the M7 Max. That’s very acceptable for an ultrabook sized laptop.

Diffusion of local LLM capable hardware throughout society

Lots of ink both physically and virtually has been spilled lamenting how Europe isn’t keeping up with the US and China on AI advancement: we aren’t investing in the electricity supply for datacentres, nor in AI research past a small fraction of what the Americans and especially the Chinese are doing. It is therefore claimed that Europe will be left behind, and left at a significant disadvantage to the US and China.

This kind of claim has been made many times before on many topics of industrial, social and political comparison between the three superpowers – and it is true that especially recently Europe has felt on the back foot as it gets bullied simultaneously by the other two superpowers, which it isn’t used to historically. However, something less appreciated is that Europe is surprisingly good at diffusing more quickly and completely the gains of an advancement than the other two superpowers: it ‘buys in’ the advancement cheap, then mass disseminates it.

That will need explaining, so to simplify: Europe, due to its unique configuration of highly competitive constituent arms length states with huge size variations, tends to diffuse innovations faster and more broadly than America or China does. This is surprising on first inspection, but think of it this way: if Ireland obtains a large current account surplus by diffusing US sourced innovations widely across its economy, all cash strapped countries elsewhere in Europe start paying rapt attention and will try to duplicate and/or improve upon whatever Ireland is doing. Ireland gets a lot of stick internationally for being a tax haven and washing the profits of US multinationals of their tax obligations elsewhere – all of which is fair – but less appreciated is that all those US multinational operated subsidiaries in Ireland do genuinely diffuse US innovations throughout the Irish economy much quicker and more completely than they could in the US where they are nowhere near as relatively economically dominant. Same goes in Switzerland and Belgium incidentally.

Obviously I’m exaggerating a touch there – at times I do wonder about diffusion of best practices in Ireland – but my point is that in superpowers such as the US and China, practice of best practices tends to be concentrated in specific economic clusters such as New York or San Francisco-San Jose in the US, or Shenzhen-Guangzhou or Shanghai in China. Whereas Europe’s economic clusters are more geographically distributed and numerous in a unique three spoke configuration:

These are the famous blue, golden and green ‘bananas’ of European economic cluster (source). Unusually they all connect together through the North of Italy, which is exactly why while the Covid pandemic may have originated in China, it turned into a global pandemic in the North of Italy as that is the most connected place to other places in world bar none other, so all global pandemics will always spread worldwide from there. As with infectious diseases, so does the global diffusion and spread of new ideas and best practices all originate from Northern Italy.

And the same will undoubtedly apply to the mass adoption and use of LLMs: the US may design the hardware and the Chinese may manufacture the hardware, but it’ll be Europe who reaps the most economic value for the cheapest price from their inventions. This is why Europe always appears to be an economic laggard, yet by all metrics it has the best quality of life for the most people anywhere in the world despite having the lowest debt to GDP ratio of any of the world superpowers. Before some say ‘that’s because you don’t spend enough on defence’, I already debunked that in past posts here: Europe has rarely spent less than the US on a PPP adjusted basis, and last few years it is by far the biggest military spender in the world (and if you include Russia in Europe, which most would, then Europe has by far and away always spent more on its military than anywhere else in PPP terms). So, in terms of economic and welfare achievement, Europe’s practice of cheaply reaping from what others sow has served it very well.

How will this affect individual behaviours and mentality?

What will the world be like when your laptop and increasingly your phone locally runs a LLM as powerful or more powerful than the world’s currently most powerful LLM?

You might think what is different to the laptop or phone using a LLM running in a cloud elsewhere and using it over a data connection?, and in some ways you would be right: I’m using a Deepseek running in some cloud elsewhere over a network connection. What’s the difference between that and running it locally?

The first difference is privacy: I wouldn’t ever put anything potentially confidential anywhere near a public internet connection. I definitely wouldn’t put any personal emails or family photos near a public internet connection. Most people won’t care, so maybe this difference only matters to people like me. Still, I’m also an individual, and for me this matters a lot.

The second difference is cost centring: if a cloud runs the LLM, somebody has to pay for that and your average individual is highly adverse to subscriptions when a free of cost substitute is available. So 98% of individuals right now use the free LLM services, and they are generally terrible because otherwise they’d cost real money. If the LLM runs on your device, you take a hit to battery life, but otherwise it’s free of cost. So for your typical individual, from 2028 onwards they’re going to experience an enormous leap in LLM capability, as until then all they’ll be used to is the crappy cheap to operate free LLMs.

The third difference is that most businesses – and a fair few individuals – don’t like to introduce single points of failure to their operations. Most cloud services are seen by many as exceedingly annoying when they go down. And the more you depend on such a service, the more anxious you get if it could disappear/get cut off/drop out. If LLMs run exclusively on hardware you personally own and control, a lot of that anxiety lifts. Now you can lock yourself into this new technology with a certainty you couldn’t have had before. It is for this exact reason why private automobiles are so popular: you aren’t buying transport from A to B, rather you’re buying the guarantee of transport from A to B which public transport only offers in big cities.

The fourth difference is that if they’re truly free of cost and you can run them all day long and all it costs you is electricity, you’re going to use them a LOT more. For everything in fact. Why search the web if your local LLM can do it for you? Why order anything or reply to any message if your local LLM can do it for you? Why think about interacting with your device if your local LLM can do it for you?

And now we’re getting into the interesting stuff: what can a LLM automate away, and what can’t it do i.e. what role is left for humans?

We don’t still know how intelligent LLMs will become before the bubble pops, but I can say this: the LLM knows more about everything than you do, but not more about some specific topics than you do. Accepting on what topics you are weak but being honest about where you genuinely really do understand more than the LLM will be the key to your success going forth.

LLMs genuinely can be a force multiplier if you use them where their strengths lie, and combine that with your strengths. But they also hallucinate and are currently lousy at direction and strategy, so that’s where I would expect the value of humans to remain. In other words, I think politicians are going to have some of the best job security going forth, because their whole purpose is to set unpopular directions for everybody else.

How will this affect individual employment?

The future world of human employment I suspect is (a) those physical jobs which can’t economically be replaced by a robot controlled by a LLM and (b) those jobs where a human’s deep understanding of a niche topic of value cannot be surpassed by any LLM, or where decisions must be taken which involve long term direction and strategy. For everything else, I expect LLMs to gradually replace all before them.

Speaking of LLM controlled robots, I was quite surprised to discover that they only melded a LLM with a humanoid robot last June, so we’ve got a few years to go before humanoid robots start taking human physical jobs. But not as many years as you might think!

Much also to my surprise, it turns out that nobody was mass producing non-toy humanoid robots until only November last year! Absolutely before then as now you can buy toy humanoid robots, these will dance for you and do kung fu etc, but they’re absolutely useless for getting any real work done as they (a) can’t lift enough weight reliably and safely and (b) they don’t have the sensors for fine dexterity manipulation in unfamiliar situations. And absolutely before now there were intelligent mass produced industrial robots – any modern factory is stuffed with them – but none were humanoid until last November.

The first mass produced industrial humanoid robot was the Ubtech Walker S2 which you can see to the right, and it went on sale in November 2025 and has probably sold about three thousand units. It costs about €150k ex VAT, it can carry up to 15 kg and you get about 2.5 hours per battery charge, though it can swap out its battery at a battery recharge station on its own so it can work continuously without a break. Its intended use is within pristine environments where fairly fixed programming works well e.g. walk over there, pick up one of X, rotate it until it has the right orientation, walk back here, put it into the right component box. In other words, just like any other industrial robot, but this one is capable of adapting to different locations within the same factory.

Next up is the Boston Dynamics Atlas which entered mass production in January 2026, and probably about two thousand units have been sold so far. It costs about €200k ex VAT, it can carry up to 30 kg including an impressive 20 kg if on one arm, and you get about two hours per battery charge. It seems a bit more intelligent than the Walker S2, but not by much: it is also intended for fixed, repetitive, work in a pristine environment like a factory floor. This robot is undoubtedly a lot more impressive in the build quality sense than the Walker S2, but it does cost a third more, and also Boston Dynamics only put it into mass production now after decades of development because they had to due to Chinese competition – not because it was finished or particularly compelling or priced well. It also is less interesting because all its production for the next two years is already sold to Hyundai, so nobody else will be able to buy one for several years more yet.

Last April, a much more interesting industrial humanoid robot went into mass production: the Figure 03. Here is a youtube live stream of it unpacking parcels in a mail office, ensuring that the address label points downwards for scanning:

Firstly, the handling of irregularly sized, sometimes squishy, items is FAR harder than the regularly sized boxes with grab handles that the previous two robots can handle. The Figure 03 can also climb stairs by itself – albeit slower than a very elderly person – but it does get there. It currently costs about €100k ex VAT, it can carry up to 20 kg and you get a very good five hours of battery life, but at the cost of it being 40% slower at movement e.g. it walks slower, moves slower etc. They have sold maybe four thousand of these by now. It comes with a bundled LLM running locally which isn’t particularly good – nowhere near even Deepseek v4 Flash in fluid conversation – but if you tell it to go wash the clothes in the washing basket it’ll go fetch the basket and take it to the washing machine, very slowly pick each item out and put it into the machine, then very slowly pour in detergent and set the washing machine running. Ultimately, apart from the thing getting in the way a lot due to its lethargy, it is a far more interesting humanoid robot – at least for the very wealthy, not least due to its cost, but also because you would really need a home with large open spaces so you can easily get around the robot while it very slowly does things. Just to be clear: the Figure 03 can run and jog as fast as a human, but it absolutely horses through its battery if it does, plus it gets hot – very hot! Still, maybe future firmware revisions could let it exchange battery for speed for short bursts so it isn’t annoying and doesn’t get in the way, but otherwise conserve battery life.

If I am being honest though, I suspect the Figure 04 is the one to wait for, as the Figure 03 feels like it has too many design and hardware compromises, and it is still too expensive for what you get on the software side. If it’s the most impressive in mass production right now, what screams out loudly is just how immature and unfinished its software story is. It’ll be years, at best, before that can be remedied.

In case you’re wondering what about all the other mass produced industrial humanoid robots, that’s it: everything else isn’t actually in mass production. In particular, Tesla’s very long advertised robot is nowhere to be seen: we don’t know its specs, its price, or anything else about it, and given Elon Musk’s long history of made up claims about autonomous driving, I wouldn’t be optimistic that his robot will have good autonomy for at least several years after launch. Ultimately this is because it’s one hard thing to build the hardware for an affordable price, it’s another hard thing to create compelling software for that hardware platform. As an example, Meta solved building affordable VR headset hardware, but they did not solve building a compelling software ecosystem for it, so the whole thing has gone off to die and it’s only a matter of time before that entire ecosystem is abandoned. Similarly, Tesla’s fully autonomous driving will likely never get solved well enough to be allowed by regulators at a price consumers will pay.

As much as Figure 03 is impressive, it still requires a pristine environment i.e. you can’t be taking it onto a building site. Even if a robot could navigate well such an irregular environment, and it coped well with getting mud and sand into its joints, it would almost certainly move too slowly for many tasks on a building site AND generally annoy the human construction workers by getting in the way.

Currently a construction worker might cost about €100k to the employer, so maybe for €100k a construction site robot might be worth the expense if it only did things like fetch concrete blocks so the blocklayers could keep working without pause. But as each block weighs 30 kg, it would need to be able to move wheelbarrows of them over scaffolding, which is far beyond the capabilities of any current or near future humanoid robot. You’d also need several of them as they’d go much slower than humans, and because they’d run out of batteries after a few hours you’d either need a quick battery swap facility or even more robots. And finally most construction sites don’t have electricity apart from a generator, so charging robots at a site would be very unattractive. So, for certainly the next decade, I think construction workers can rest in peace that they will not get made unemployed by humanoid robots.

For human jobs doing physical labour in pristine environments though, the next ten years looks like increasing levels of human jobs being displaced. If they can get the cost of these robots down to €25k, a lot of minimum wage jobs like stacking shelves or packing online orders look inevitably gone forever, as the minimum cost to the employer of a human (minimum wage is about €32k) makes the robot look cheaper. If somebody successfully cracks deep cleaning by robot, that’s all your cleaning staff gone too. Jobs like fast food kitchen work is at threat, even if the delivery driver is not – that’s a lot of your young person entry level jobs disappearing forever there.

The most recent (2024) ESRI report lists the largest number of minimum wage jobs being in these sectors:

  1. Kitchen helpers (14%)
  2. Shop sales assistants (10%)
  3. Bartenders (7%)
  4. Caretakers (6%)
  5. Waiters (6%)
  6. Home based personal care (3%)
  7. Housekeepers (2%)
  8. Receptionists (2%)

About ten percent of the Irish workforce earns near minimum wage, and humanoid robots could take over most if not all those eight sectors if they get cheap enough.

Food for thought indeed! This displacement of humans from their jobs by AI might have impacted IT first, but I am extremely sure it shall be coming for entire sectors of knowledge worker and pristine environment manual labour next.

#LLM




Go back to the archive index Go back to the latest entries

Contact the webmaster: Niall Douglas @ webmaster2<at symbol>nedprod.com (Last updated: 2026-08-07 00:01:44 +0000 UTC)