Not everyone who solves a problem gets to be remembered for solving it first.
The Thesis Nobody Opened
In 1974, a Harvard doctoral student named Paul Werbos sat down and worked out, in full mathematical detail, a way to teach a multi-layer network to learn from its own mistakes. He wasn't guessing. He wasn't close. He had it — the same core idea this week's video walks you through, more than a decade before it became famous.
Almost nobody read it. Werbos buried the neural-network application inside a dissertation mostly about economic forecasting, and he was publishing it into the same climate we covered back in Episode 8: a research world where "neural network" was close to a dirty phrase, funding had dried up, and reviewers had little patience for an approach the field had already declared a dead end. A brilliant answer, arriving in a room where the question had been ruled out of order, doesn't get much of an audience.
The Story Behind the Story
Here's the part the video doesn't have room for: Werbos wasn't even the only one. Around 1985, a Stanford researcher named David Parker worked out essentially the same backward-calculation trick, independently, with no idea Werbos's thesis existed. Across the Atlantic, a young French researcher named Yann LeCun arrived at a close cousin of the same idea, also without knowing about Werbos or Parker. Three people, three continents of academic distance, converging on the same insight — and none of them aware the others had gotten there first.
Then in 1986, David Rumelhart, Geoffrey Hinton, and Ronald Williams published the paper that finally made it stick — not because their math was more original, but because their paper made the case impossible to dismiss. Timing, it turns out, mattered as much as the discovery itself.
There's a quieter thread in this story, too. Rumelhart went on to spend the following years doing exactly what backpropagation made possible — building the theoretical foundations of how machines learn. In the 1990s, a rare neurodegenerative illness began taking that capacity from him, piece by piece, well before the deep learning boom he'd helped make possible reached the rest of the world. Colleagues established a prize in his name while he was still alive to make sure his contribution wouldn't quietly disappear the way Werbos's almost had. He died in 2011, one year before the field he shaped exploded into the mainstream.
What This Really Means
This isn't a one-off story. It's a pattern that shows up again and again in AI's history: the right idea, discovered in the wrong decade, sitting unread until the culture catches up to it. Being first has never guaranteed being remembered. What tends to get remembered is whoever publishes at the moment the field is finally ready to listen — and that moment is often decided by funding, mood, and luck as much as by math.
Watch the Mechanism
The video breaks down exactly how the calculation Werbos, Parker, LeCun, and the 1986 trio all separately stumbled onto actually works — the forward pass, the single number that measures a mistake, and the backward sweep that hands every weight in the network its precise share of the blame. Once you've read how close the world came to losing this idea entirely, watching the mechanism finally click into place hits a little differently.
Some ideas don't fail because they're wrong. They fail because nobody was ready to hear them yet.


