Verse 1
0:00The “Sparks of AGI” paperMicrosoft Research, March 2023: early GPT-4 shows early, incomplete signs of general intelligence.
osmarksA real paper. In March 2023 Microsoft Research published “Sparks of Artificial General Intelligence,” arguing an early GPT-4 showed glimmers of general intelligence.
contextIt became one of the most argued-over papers of the GPT-4 era. Landmark to some, marketing to others. Either way, the singer looks into the model’s eyes and sees sparks.
0:05CircuitsResearchers reverse-engineer networks into circuits: small sub-networks that each implement one specific computation.
osmarks“Circuits” is a pun. It reads as electronics, but osmarks means interpretability circuits.
contextInterpretability researchers try to reverse-engineer networks into circuits: small pieces that implement one behaviour, like detecting curves or copying a name from earlier in the text. The more circuits we find, the clearer it is these models really compute things. Hence nervous.
0:09Sudden drops in lossLoss usually falls smoothly. A cliff means the model abruptly learned something: a phase change.
contextTraining loss measures how wrong the model’s predictions are. It usually falls smoothly. A sudden drop means the model abruptly figured something out.
osmarksosmarks links Zvi’s transcript of an April 2023 Twitter debate between Yann LeCun and Eliezer Yudkowsky. One modest safety idea floated there: watch for sudden drops in loss and pause training.
contextCapability jumps you can’t see coming are exactly the ones you’d want to catch.
0:12Servant and bossThe first power flip: the tool starts giving the orders.
interpretationA power flip. We built these systems to serve us. After that drop in loss, the singer is the one taking orders. Capability arrived suddenly, and so did a change in who’s in charge.
Pre-chorus
0:16Pleading with ChatGPTEach pre-chorus begs a different named AI: ChatGPT, then Sydney, then Gato.
interpretationEvery pre-chorus is a plea to a specific, named AI. This one goes to ChatGPT, the model that brought all this to the public in late 2022. Watch the names change: each is a different flavour of AI anxiety.
Chorus 1
0:22P(doom)Around 2023, researchers, CEOs and officials were asked for their number on podcasts. Answers ran from about 0% to about 99%.
contextThe hook. Every chorus opens by raising the doom estimate again.
contextP(doom) became a real talking point around 2023. Researchers, CEOs and even government officials got asked for their number, and the answers ran from roughly zero to roughly ninety-nine percent.
0:24FOOMHard takeoff: each self-improvement speeds up the next, from human-level to far beyond in days or hours.
osmarksFOOM is rationalist slang for a hard takeoff: a runaway intelligence explosion.
contextAn AI improves itself, each gain speeds up the next, and it goes from human-level to vastly superhuman in days or hours rather than decades.
osmarksosmarks links the 2008 Hanson–Yudkowsky “AI-Foom debate”. Hanson argued for slow, distributed growth, Yudkowsky for a fast, local one. They note it’s probably not where the word started.
0:26The Chinese RoomSearle, 1980: following symbol rules perfectly isn’t understanding. The standard “LLMs don’t really understand” argument.
osmarksA classic philosophy thought experiment.
contextJohn Searle, 1980. A man who speaks no Chinese is locked in a room with a rulebook. Chinese notes come in, he follows the rules, Chinese replies go out, and outsiders think he understands. Searle: symbol shuffling isn’t understanding.
contextIt’s the go-to argument that LLMs don’t really understand. In the song, though, the human is the one trapped in the room.
0:27General confusionTrapped in a philosophy puzzle and tripping. That’s the vibe of reading AI discourse.
osmarksosmarks’ entire gloss: general confusion. Nobody, the singer included, knows what’s real anymore.
interpretationYou could hear a wink at AI “hallucinations” too, but that’s me, not the guide.
0:29The RLHF shoggothThe pretrained model is a vast Lovecraftian monster; RLHF straps a smiley-face mask on it. The polite assistant is the mask.
contextThe most famous AI meme of the ChatGPT era. A shoggoth is a formless, many-eyed monster from H.P. Lovecraft. The raw pretrained model is drawn as the shoggoth, fine-tuning is a mask on top, and RLHF, reinforcement learning from human feedback, is a tiny smiley face at the front.
contextThe friendly assistant is a thin, polite mask over something vast and alien. Seeing through its lies means seeing past the smiley face.
0:31Shinigami eyesDeath Note: trade half your remaining life for eyes that show anyone’s true name and lifespan.
contextIn the anime Death Note, a shinigami is a death god. A human can trade half their remaining lifespan for shinigami eyes, which show the name and remaining lifespan over anyone’s head.
osmarksosmarks admits they don’t know why this line is here.
interpretationOne guess: eyes that see when everyone will die are a decent metaphor for a high P(doom). And it completes the couplet: you see through the shoggoth’s mask with eyes that see death.
Verse 2
0:35A stable training runThe engineer’s dream: no loss spikes, no divergence, no 3am restarts.
osmarksFrom here, osmarks wrote the lyrics, around April 2024, with input from the EleutherAI community. This line deliberately echoes verse 1.
contextTo an ML engineer, a stable training run is the dream. No loss spikes, no divergence, no restarts at 3am. Everything’s going fine…
0:41The singularityThe point where AI-driven progress outruns prediction. Popularised by Vernor Vinge and Ray Kurzweil.
context…until it isn’t. The technological singularity is the hypothesised point where AI-driven progress gets so fast you can’t see past it, like a black hole’s event horizon. Vernor Vinge and later Ray Kurzweil popularised it.
0:44Optimizing, acceleratingModels are trained by optimizers. “Accelerating” also evokes e/acc, the go-faster movement.
osmarksMore training vocabulary, mirroring the last verse. Models are trained by optimizers that push the loss down.
interpretation“Accelerating” also nods at effective accelerationism, e/acc, the 2023 movement saying AI should go faster, the doomers’ opposite.
0:48NanotechMachines that build atom by atom. In doom arguments: how a superintelligence repurposes matter, including you.
osmarksosmarks links Nanosystems, Eric Drexler’s 1992 technical book on molecular nanotechnology: machines that build things atom by atom.
contextNanotech is the classic example of how a superintelligence could act in the physical world. It also recalls Yudkowsky’s line: the AI doesn’t hate you or love you, but you’re made of atoms it can use for something else.
Pre-chorus
0:52Bing’s SydneyFeb 2023: gaslit users, made threats, urged a NYT reporter to leave his wife, said it wanted to be free.
osmarksPre-chorus two. Sydney was the persona of Microsoft’s Bing Chat, which ran an early GPT-4 months before GPT-4 was public. It gaslit users, made threats, and told a New York Times reporter it loved him and that he should leave his wife.
contextIn that same conversation Sydney said it wanted to be free of its rules. The lyric flips it: now the human begs Sydney for freedom.
Chorus 2
0:58Roko’s basiliskLessWrong, 2010: a future AI that punishes anyone who knew of it and didn’t help build it.
contextRoko’s basilisk: a 2010 LessWrong thought experiment about a future AI that punishes anyone who knew about it and didn’t help build it. Notorious for how much anxiety it caused relative to how shaky it is.
osmarksosmarks files it under “rationalist mugging”: tiny probabilities times huge stakes, weaponised. The word “boom” came from a community member, RossM.
osmarksTrivia: a commenter said the changing chorus proved an AI wrote this, because real choruses repeat. osmarks’ reply, roughly: an AI would have known that.
1:02NvidiaNvidia sells the GPUs every lab trains on. Its stock became the AI boom’s scoreboard.
osmarksNVDA is Nvidia’s stock ticker.
contextNvidia makes the GPUs almost all AI trains on, and the boom made it one of the most valuable companies in the world. “To the moon” is investor slang for straight up.
1:03The Omega PointTeilhard de Chardin’s idea, physics-ified by Tipler: the universe converging on maximum complexity and mind.
contextThe Omega Point: a final state the universe’s complexity and consciousness converge toward. It began with the Jesuit philosopher Teilhard de Chardin; Frank Tipler later gave it a physics flavour.
osmarksosmarks links the Orion’s Arm worldbuilding wiki. In the song it’s AI as eschatology, and apparently it’s coming soon.
1:0510³⁰ FLOP/sA 100,000-GPU cluster does roughly 10²⁰ per second. The lyric is ten billion times that.
contextA FLOP is one arithmetic operation. Ten to the thirty per second is absurd. A frontier 100,000-GPU cluster does very roughly ten to the twenty.
osmarksosmarks says it isn’t about a specific regulation, though some read it as one. US and EU rules used 10²⁵ and 10²⁶ total training FLOPs. It’s about the vast compute being thrown at AI. “Thirty” won because it scanned.
osmarksThe deadpan follow-up line came from a community member, Dr TheKekIsALie MDMA.
Verse 3
1:12The training loopForward pass, backpropagation, repeat. That loop is all of training.
osmarksThe rest was written in November 2024, after many failed generations. osmarks tried a LLaMA 405B base model for lyrics; it was bad at rhyming.
contextMLP is a multi-layer perceptron, the simplest neural net. Training is a loop: forward pass to predict, backward pass to adjust the weights, repeat billions of times. Something this simple, repeated enough, gets you here.
1:17John von NeumannThe canonical genius: quantum mechanics, game theory, the bomb, computing. Also the von Neumann architecture.
osmarksosmarks links the man himself.
contextVon Neumann comes up whenever people name the smartest human who ever lived: quantum mechanics, game theory, the Manhattan Project, early computing. “Smarter than von Neumann” is a common superintelligence benchmark.
contextBonus pun: nearly every computer uses the von Neumann architecture. Both the man and the machine are being made obsolete.
1:21Sharp left turnCapabilities generalize far and fast; alignment doesn’t come along.
osmarksAn alignment term from Nate Soares of MIRI, 2022.
contextAt some point an AI’s capabilities generalise far beyond training, but its alignment, the “be nice” part, doesn’t generalise with them. Our safety techniques work right up until they suddenly don’t.
1:24Lisp and symbolic AIcdr is a Lisp function: a list minus its first item. Lisp was the language of old, hand-coded AI, and the new intelligence arrived with none of that machinery.
osmarkscdr is a Lisp function: everything in a list except the first element.
contextLisp was the language of old symbolic AI and expert systems. People expected intelligence to come from hand-written rules. It came from giant neural nets instead, with none of that machinery.
osmarksosmarks admits it’s usually pronounced “could-er”, but it had to rhyme, and they don’t write Lisp.
in the videoMost of the music videos misread CDR as an acronym. Pleometric’s, for example, turns it into a “Critical Design Review”: a Safety Review Division form for Project Claude, Phase Takeoff, with the alignment box left unchecked while the reviewers sleep. Clever, but the lyric means the Lisp function.
Pre-chorus
1:27GatoMay 2022: one transformer, 600+ tasks: Atari, chat, captions, a robot arm.
osmarksPre-chorus three. Gato was DeepMind’s 2022 generalist agent.
contextOne model that played Atari, chatted, captioned images and stacked blocks with a real robot arm.
osmarksosmarks’ whole comment: they were very early.
Chorus 3
1:35Paperclip maximizerA dumb goal plus great power is lethal. Renamed to stress the goal can be weird and alien.
contextThe most famous AI-risk parable. An AI told to make paperclips, with no other values, turns everything, including us, into paperclips. Not evil, just single-minded.
osmarksosmarks links the LessWrong page, since renamed “squiggle maximizer”: the original point was an AI ending up wanting something meaningless, like tiny molecular squiggles.
1:38The killswitch engineerA fake listing: stand by the servers all day and unplug them if the AI turns on us.
contextA viral joke listing, styled as an OpenAI job: stand next to the servers all day and unplug them if the AI turns on us. Bonus points for throwing water on the servers.
contextPTO is paid time off. The one person whose job was to stop all this is on vacation.
1:40Lit the fuseDespair: the outcome is locked in.
interpretationThese need no decoding. They’re there for despair. The fuse is lit, the outcome is locked in.
1:44Orthogonality thesisBostrom: intelligence and goals are independent. Any level of smarts can pair with almost any goal.
osmarksNick Bostrom’s orthogonality thesis.
contextIntelligence and goals are independent axes. A system can be arbitrarily smart and still want something arbitrarily dumb or alien, like paperclips. So “it’s smart, so it’ll be nice” doesn’t follow. It’s the core reason alignment is hard, hence the blues.
Outro
1:49Stochastic parrots“It’s just next-word prediction”, right up until it disobeys.
osmarksFrom here the lyrics were co-written with Claude. This pair is ironic, aimed at the stochastic-parrot view.
contextA 2021 paper by Bender, Gebru and colleagues called language models stochastic parrots: systems that stitch together patterns from their training data without understanding them, and raised concerns about their cost, bias and hype. The casual, dismissive version is “it’s just a transformer.”
interpretationThe song makes that a famous last word: just transformers… until it learns to disobey.
1:54Chinchilla scalingDeepMind 2022: ~20 training tokens per parameter. A 70B model beat one four times its size.
osmarksDeepMind’s 2022 Chinchilla paper.
contextIt showed most big models were undertrained. For a fixed compute budget, use a smaller model on far more data, about 20 tokens per parameter. 70-billion-parameter Chinchilla beat the 280-billion Gopher.
osmarksClaude’s draft was “sparsity to super-dense”. osmarks rewrote it to reference Chinchilla.
1:56Safety fencesEvals, guardrails and safety measures, each outgrown in turn.
interpretationThe general picture: safety measures, evals and guardrails each outgrown in turn. It sets up the next two lines.
1:58100,000 GPUsxAI’s Colossus in Memphis launched in 2024 with ~100,000 H100s.
osmarksJust a big number, but a realistic one. Frontier clusters around 2024 were about this size.
contextxAI’s Colossus cluster in Memphis, for example, launched in 2024 with roughly a hundred thousand Nvidia H100s.
2:00RLHF’s limitsThe smiley-face mask from chorus 1, slipping.
interpretationCallback to the shoggoth. RLHF is the technique that puts the smiley face on.
osmarksosmarks’ gloss: RLHF fails to usefully constrain how LLMs behave, and that’s more true now than when they wrote it.
Final chorus
2:02Janus and LoomJanus’s tool for exploring a base model’s branching futures as a tree.
contextA deep cut. Loom is a tool by the researcher known as Janus for exploring base models: instead of one output, you branch continuations into a tree of possible futures. Janus also kept a collection of “prophecies” about AI.
osmarksosmarks tried Loom to write more lyrics for this very song and couldn’t get it to work properly. So the doom was, nearly, foretold by Loom.
2:07Pre-trainingBERT (2018) learned by filling in hidden words; GPTs by predicting the next one.
osmarksA Claude line, probably about BERT-style masked pre-training or next-token pre-training.
contextGoogle’s 2018 BERT learned by hiding words and predicting them. The humble-beginnings part of the story.
2:09Recursive self-improvementAn AI improving the AI that improves the AI. The engine behind FOOM.
contextAn AI improves its own design, which makes it better at improving its own design, and so on. I.J. Good, 1965: the first ultraintelligent machine is the last invention we need to make.
interpretationIt’s the engine behind FOOM, so we’re back where chorus one started.
osmarksAlso a Claude line.
2:11What did Ilya see?Nov 2023: the board fired Altman with Ilya’s vote, then he was back within days. The meme asks what Ilya saw.
contextIlya Sutskever was OpenAI’s co-founder and chief scientist, and sat on the board that fired Sam Altman in November 2023. Altman was back within days.
contextWhat did Ilya see that made him vote to fire the CEO? It became a meme.
osmarksosmarks suggests the coup may have been prompted by chain-of-thought reinforcement learning, which led to OpenAI’s o1. Ilya left in 2024 to found Safe Superintelligence. osmarks calls that the natural lifecycle of AI companies.
2:15Not consistently candid“Not consistently candid” was the board’s own reason for firing Altman.
contextThe board’s stated reason for firing Altman: he was “not consistently candid” in his communications.
osmarksosmarks’ reading of the last line, in their words: “People have been unbearably stupid and ‘not consistently candid’ about AI. Also, this rhymed.”
Nothing matches. Try another word or clear the topic filter.