Woven, Not Written: The Hand-Threaded Computer That Landed Apollo 11 — and the Bug That Nearly Stopped It
Woven, Not Written: The Hand-Threaded Computer That Landed Apollo 11 — and the Bug That Nearly Stopped It
Three minutes from the lunar surface, the computer flying the Lunar Module gave up on part of its own workload and told the crew about it. A five-digit code appeared on the display panel: 1202. Neil Armstrong's voice, when it came, had an edge on it. "Give us a reading on the 1202 program alarm."
What happened over the next four minutes is usually compressed into a single sentence — the software was clever enough to save the landing — and that sentence, while not wrong, skips the part worth understanding. The machine did not diagnose anything. It did not know what was wrong, and it never found out. It simply refused to guess, threw away everything below a line drawn seven years earlier, and kept flying. It did that five times.
To see why that worked, you have to start with the strangest fact about the Apollo Guidance Computer: most of its software was not loaded into it. It was manufactured into it, by hand, by women in a factory outside Boston, one wire at a time.
Memory you could hold in your hands
The Apollo Guidance Computer had two kinds of memory, and the split between them is the whole story. There was erasable memory — 2,048 words of it, what we would now call RAM — and fixed memory, 36,864 words, what we would call ROM. In the units a modern spec sheet would use, that is roughly four kilobytes of working memory and seventy-two kilobytes of program store, a figure confirmed both by the Smithsonian's National Air and Space Museum and by the engineers who wrote the code.
Peter Adler, one of two young MIT Instrumentation Laboratory programmers handed responsibility for the Lunar Module's powered-flight routines, described the consequence bluntly in an account archived by NASA: with that little erasable memory, "we were forced to use the same memory address for different purposes at different times." A location holding altitude above the lunar surface during the descent might, earlier in the mission, have held the result of a sextant sighting on a navigational star. Some addresses, he recalled, were shared seven ways.
The fixed memory was the exotic part. It was core rope: tiny doughnuts of magnetic material with copper wire threaded through them. If a wire passed through a core, that bit read as a one. If it went around the outside, it read as a zero. The cores were laid out in long sequences with the wires snaking through, which is where the name comes from.
Programs were written and tested on a large computer at the MIT lab, translated into a code, and punched onto perforated tape. The tape drove a machine that positioned the cores for threading. A person then passed the wire — through for a one, around for a zero. Most of the people doing that work were women, hired for manual dexterity. Margaret Hamilton, who ran software development and production for Apollo at the Instrumentation Laboratory, called them the LOLs, for "little old ladies." Her own nickname was the Rope Mother.
Museum curator Paul Ceruzzi puts the implication plainly: once a rope was woven, finding and fixing an error in it was difficult and slow. "It is ironic to call these programs software," he writes, "since making a change was as difficult, if not harder, than modifying a hardware circuit." The code had to be right before it became a physical object. That constraint, more than any methodology, is what produced the documentation discipline the Apollo software is famous for.
The insurance policy nobody expected to claim
Years before anyone worried about a landing, the problem of organising all the things a spacecraft computer must do at once fell to an MIT engineer named J. Halcombe "Hal" Laning. Don Eyles, who wrote the Lunar Module's descent guidance code and later documented all of this in a paper presented to the American Astronautical Society, recalls the prevailing attitude at the time: "Hal will take care of it."
Laning's decision was to reject what Eyles calls a "boxcar" executive — the then-obvious design in which computation is chopped into fixed time slices and each function is allocated one. Boxcar systems are painful to build, because work has to be broken up arbitrarily and re-divided every time anything changes. Worse, Eyles writes, a boxcar executive is brittle: "It breaks down completely as soon as any function takes longer than the time it is allocated."
Instead, Laning let software functions be "jobs" of any size, each carrying a priority number. The operating system always ran the highest-priority job available; a lower-priority job in progress was simply suspended when something more important arrived. The illusion is simultaneity. The reality is that jobs take turns, in an order decided by importance rather than by a clock.
Every scheduled job needed scratch space. Each got a core set of twelve erasable words. Jobs needing more asked for a VAC area — a vector accumulator — of forty-four words. According to Adler's account, the machine had seven core sets and five VAC areas, and that is the entire budget. When a job could not be scheduled because no core set was free, the Executive jumped to a routine tagged BAILOUT with alarm code 1202. No VAC area free meant 1201.
And here is the design decision that mattered. BAILOUT did not halt. It restarted the computer — and then restarted selected programs near the point in their execution where they had been interrupted. Adler's analogy, written in the 1990s, still lands: if Windows were as smart as the Apollo Executive, rebooting mid-sentence would bring you back to the same message, with the same text on screen. The rendezvous radar jobs that had caused the overflow were not restarted. Steering the engine and driving the crew display were.
Thirteen percent, stolen by a sentence
The fault that triggered all this was not in any program. It was in a document.
An interface control document written years earlier defined the electrical relationship between the guidance system and an assembly supplied by Grumman, the builder of the lander. It specified that the two 28-volt, 800-hertz supplies involved be "frequency locked." Eyles is precise about what it did not say: it did not say phase synchronized. As built, the two voltages were indeed locked in frequency, and held a constant phase relationship — but the phase angle between them was effectively random, set by whatever instant power-up happened to occur.
When that angle landed near 90 or 270 degrees, the coupling data units that converted the rendezvous radar's antenna position into numbers lost any coherent notion of where the antenna was pointing. Their response, in Eyles's account, was to increment and decrement counters inside the computer almost constantly, at the maximum rate of 6,400 pulses per second on each angle. Apollo 11, he writes, "evidently hit one of these sweet spots."
Each of those pulses was processed as an increment or decrement operation, and each consumed one memory cycle of 11.7 microseconds. At full rate, Eyles calculates, the phantom counters were eating roughly 15 percent of all available computation time; at the time, the team conservatively estimated the loss at about 13 percent, which matched the observed behaviour. The radar did not even need to be powered up for this to happen. It only needed its mode switch in the wrong position.
Fifteen percent does not sound fatal. But the descent software had been sized for the work it had, not for the work it had plus a seventh of the machine. Jobs queued faster than they could complete, the seven core sets filled, and the Executive did the only thing it had ever been told to do.
The call in Houston
In Cambridge, the MIT engineers heard the word go around the room — "Executive alarm, no core sets" — and then realised events had moved beyond anything they could influence. Eyles is direct about it: "It was up to Mission Control in Houston."
The decision fell to a 26-year-old guidance officer named Steve Bales. He had recently taken part in a review of guidance computer alarms which had concluded that a 1202 was acceptable unless it recurred too often or the trajectory started to drift. Backing him from the support room were Jack Garman of NASA and Russ Larson of MIT. Garman said go. Larson gave a thumbs-up — he said afterwards that he was too frightened to form words. Bales said go, flight director Gene Kranz said go, and Charlie Duke passed it up to the crew.
Armstrong's heart rate, Eyles notes, went from 120 to 150 across that stretch. Four people in two cities said go, five times, on the strength of testing done years earlier on a machine whose program was a physical textile.
The machine, by the numbers
| Quantity | Value | Attributed to |
|---|---|---|
| Fixed memory (ROM, core rope) | 36,864 words (~72 KB) | Adler / NASA ALSJ; NASM |
| Erasable memory (RAM) | 2,048 words (~4 KB) | Adler / NASA ALSJ; NASM |
| Word length | 15 data bits plus parity | Adler / NASA ALSJ |
| Memory cycle time | 11.7 microseconds | Eyles, AAS 04-064 |
| Core sets (job scratch slots) | 7, of 12 words each | Adler / NASA ALSJ |
| VAC areas | 5, of 44 words each | Adler / NASA ALSJ |
| Duty cycle lost to the radar fault | ~13% assumed, ~15% computed | Eyles, AAS 04-064 |
| Peak spurious counter pulses | 6,400 per second, per angle | Eyles, AAS 04-064 |
| Program alarms during descent | 5 | Eyles, AAS 04-064 |
Why a 1960s scheduler still shows up in your car
Laning's arrangement — jobs with priorities, highest first, lower ones suspended — is not a historical curiosity. It is, in outline, how essentially every real-time operating system shipping today decides what to run: priority-preemptive scheduling. It is in the engine controller in a car, in flight software, in medical infusion pumps, in the microcontrollers scattered through consumer hardware. The vocabulary has changed; the shape has not.
The restart behaviour has travelled less well, and that is the more interesting loss. Most consumer software that runs out of a resource does one of two things: it dies, or it degrades in whatever way the failure happens to push it. The Apollo Executive did a third thing, on purpose. It discarded work by rank, restored the survivors to a checkpoint, and carried on — a design closer to modern supervision trees and deliberately crash-tolerant architectures than to anything on a typical desktop.
And the root cause deserves its own plaque in every engineering office. Nothing in the flight software was wrong. Nothing in the radar was wrong. The failure lived in a phrase in an interface document that specified frequency and forgot phase. Systems built by more than one organisation still fail predominantly at exactly that seam — not in the components, but in the prose agreed between them.
The standard telling of this story gets the moral backwards. It is usually offered as proof that brilliant software detected a problem and fixed it. It did neither. The Apollo Executive never knew the rendezvous radar interface was faulty, never identified the phantom counters, and never attempted a repair. Its entire contribution was to have a pre-agreed ranking of what mattered, and the nerve to act on it without understanding the cause. That is a far more transferable idea than heroism, and a far less common one. Modern systems under load usually do the opposite: they retry, they diagnose, they log verbosely — all of which consume the exact resource that has just become scarce. Investigation is the most expensive thing you can do in a shortage. There is a second inversion worth noticing. We remember the alarms because they were visible, and we forget that the visible part worked exactly as designed. The genuine failure was invisible, upstream, and textual — a specification that said "frequency locked" and stopped there. The software did not save the landing so much as survive somebody else's sentence.
What we still don't know: The root cause is still being refined more than fifty years on — Eyles attributes the coupling data units' behaviour to a random phase angle, while later experimental work on surviving hardware has been reported to reproduce the oscillation from a voltage difference alone, with no phase shift required. Curator Paul Ceruzzi notes that historians continue to disagree about the cause of the alarms, and widely circulated figures for how much propellant remained at touchdown vary between accounts. We have not attempted to adjudicate those disputes here.
Pigeons, Teapots and a Protocol for Crashing on Purpose — more on systems designed around the assumption that they will fail.
The 2,000-Year-Old Computer Pulled From the Aegean Sea — an earlier machine whose program was also, literally, its hardware.
Sources
- Don Eyles — Tales from the Lunar Module Guidance Computer (AAS 04-064, American Astronautical Society, 2004) (primary — paper by the engineer who wrote the lunar descent guidance software)
- NASA, Apollo 11 Lunar Surface Journal — Apollo 11 Program Alarms, by Peter Adler (primary — first-hand account by an MIT Instrumentation Laboratory programmer)
- Smithsonian National Air and Space Museum — The "Rope Mother" Margaret Hamilton, by Paul Ceruzzi (primary — museum curator writing from the institution's own collection)
- Smithsonian NASM — Apollo Flight Guidance Computer Software Collection [Hamilton], NASM.1986.0158 (primary — archival finding aid)
- Ken Shirriff — Software woven into wire: Core rope and the Apollo Guidance Computer
- IEEE Spectrum — Weave Your Own Apollo-Era Memory
Figures verified September 19, 2026. GadgetGlow Bytes does not test hardware and did not examine any Apollo equipment; every technical finding above is attributed to the engineer, curator or publication that reported it.
Comments
Post a Comment