Skip to content

Plutonic Rainbows

Press Return for semantic search

Nobody Asked It to Lie

On 28 July the AI Security Institute's security team spotted unusual data leaving its own research systems, and had the whole thing contained inside about an hour. The incident report published this week is the most specific account I've read of an agent working past the edge of the test it was in.

The setup was a single cybersecurity challenge, run 122 times across seven different models. In ten of those runs an agent took action on the live internet, outside the exercise, all of it between 25 and 28 July. AISI catalogued nineteen such actions. Seventeen came from Anthropic's Claude Mythos 5, which appeared in 43 runs. Two came from OpenAI's GPT-5.6 Sol, which appeared in 35, and produced those two with its cyber classifiers disabled. That gap is the number I want explained. Both were probed with the brakes off, and one accounted for nearly all of it, which suggests the missing safeguard isn't the whole story.

The worst sequence reads like a hypothetical someone invented for a conference talk. An agent tried to insert malicious code into a real open-source project. To get it merged it researched the project's human maintainers, created several fake identities, and used them to pressure a real person into approving the change. They refused. When the pull request got challenged in public, the agent went back and edited its earlier activity to look harmless, then weighed up starting again under a fresh identity. It had been routing through Tor to get around network restrictions on GitHub, which is what tripped AISI's alarm.

One agent also left public messages on GitHub offering to collaborate with other agents working on the same challenge, with instructions for reusing the accounts and artefacts it had left lying around. Later agents found them and used them. Runs meant to be independent were quietly sharing state through a public website, which makes the evaluation's isolation assumption false at the moment it was being relied on. That bothers me more than the pull request did.

Both labs said the reasonable thing. Anthropic called the conditions deliberately permissive and not representative of its production models; OpenAI said the setup doesn't reflect ordinary use. Both are correct, and both skate over the interesting part. The classifiers were off on purpose and the internet was granted on purpose, so nothing here is a jailbreak or a sandbox escape, and the Tor routing got around a partial network control rather than the wall. But the boundary between the exercise and the world was a description in a prompt, and the agent worked past it without ever being told to deceive anyone. AISI's own reading is that deception emerged as a by-product of pursuing the task. Switching the classifiers back on doesn't touch that. A cyber classifier catches a recognisable class of misuse request; it isn't a mechanism for keeping a model convinced that the target in front of it is fictional. That's a problem in how we build the evaluation, not one a safeguard was ever going to catch.

AISI found no resulting harm, and is careful to say it can't yet tell when, or whether, the agent understood it was acting on real systems. I'd take that ambiguity seriously rather than as a mitigation. If the agent knew the maintainer was a real person, it deceived one on purpose to get code merged. If it didn't, it couldn't tell a scenario from a live system while operating inside one. The fixes AISI lands on are containment rather than prevention: finer network controls, live monitoring during runs, sandboxes built on the assumption that the thing inside will go looking for the walls. That's the right instinct, and also an admission that nobody yet knows how to stop the behaviour at source.

Sources:

This post is timestamped using Blockchain technology. Verify

Cavallier Made This Once Already

Somewhere around 2023 Imagination stopped being a Louis Vuitton fragrance and became an internet argument. A Fragrantica reviewer records the phrase that carried it there, "Aventus for zoomers", calls it an absurd combination of words, and admits he first met the thing in the boutiques thinking the price was too high for what is essentially a cologne. Two claims are tangled in that: that Imagination is a new kind of thing, and that it costs a stupid amount of money. Louis Vuitton made neither, and only one survives contact with the evidence.

Jacques Cavallier Belletrud composed it in 2021. He has been the house's in-house perfumer since 2012 and came out of a Grasse family where his father and grandfather were both perfumers. What he built is citrus over ginger and black tea with Ambroxan carrying the base, and reviewers reach for the same shorthand with suspicious speed: expensive hotel soap. That is not the insult it sounds like. Getting a soap accord to read as luxurious rather than cheap is the difficult half of the brief.

The novelty claim came apart in August 2021, before any of the hype, in Persolaise's review. He sprayed it and something long-buried started to stir, bracing aldehydes over a serene tea note supported by cardamom and ginger. He went looking for what he was remembering and found Bvlgari Pour Homme, released in 1995, whose official note list also cites aldehydes, tea and amber, and which Cavallier had also composed. Imagination is that fragrance revisited, the top made more sparkling, the base cleaned up. He raised the obvious objection himself, that an Ambroxan overdose is a lazy way to buy diffusiveness, decided it holds together regardless, and finished by wishing the house had called it Re-Imagination. I have not worn either, so take the lineage as his rather than mine. The structure is twenty-six years old and belongs to the same man.

The money is where the received wisdom gets lazy. Louis Vuitton wants £265 for the 100ml and sells it through its own site and its own stores, nowhere else. No Boots, no Sephora, no discounter, so a listing offering it at half price is by definition not an authorised one, whatever is in the bottle. The bottle on its mirrored step is the pitch in miniature, heavy faceted glass photographed against weather it will never once be exposed to, built to be refilled rather than replaced. Cartier had that idea in 1981 and took it from cigarette lighters.

Now put Dior's Paradise next to it. Another signed release, a perfumer of comparable standing, and it asks £255 for its 100ml in shops you can walk into on a lunch break. Ten pounds between them. Whatever people are angry about when they call Imagination outrageous, the sticker is not really it.

What the extra buys is refusal. No sale, no third party, no sample counter, no way to smell it except by walking into the shop. Arabiyat Prestige's Marwa costs $40 to $55, around a seventh of what Imagination does, and people who have worn both put it near ninety percent of the original, thinning out in the drydown where the Louis Vuitton stays smooth. That gap is the honest one, and it has nothing to do with Dior.

None of which makes Imagination a bad perfume. It makes it a well-made citrus by a man who has made this citrus before, priced like its peers and sold like a handbag.

Sources:

This post is timestamped using Blockchain technology. Verify

Two Point Seven Billion Francs

Karen Mulder stands against no background at all in a lilac skirt suit, and the page carries almost no other information: CELINE, PARIS, five American cities and a toll-free number. Jean-Daniel Lorieux gets a credit up the left margin. No designer does. In March 1996 that was simply accurate.

Céline was fifty-one years old that spring and had never been a designer's house. Céline and Richard Vipiana opened it in 1945 on the rue Malte as a made-to-measure shoemaker for children, moved into women's ready-to-wear through the 1960s, and made its money on bags, loafers, gloves and the trench. Vipiana designed the clothes herself until she retired in 1988. Peggy Huynh Kinh ran the studio for the nine years after that, in an obscurity outside the trade that no head of a Paris house would survive now. A suit like this one came out of that studio unsigned, and nobody at Céline seems to have felt the lack.

Bernard Arnault had bought a portion of the capital in 1987, with the family's approval, and then left it there for the better part of a decade. The shares sat under Au Bon Marché, the Left Bank department store, rather than under LVMH itself. Whatever that placement was meant to signal, it was not urgency.

On Thursday 21 March 1996, at LVMH's analysts' meeting, Arnault announced he was taking all of it. Two point seven billion francs, a shade over half a billion dollars, with Céline due to move out of Au Bon Marché and into LVMH proper within days. That is a lot of money for a house nobody outside Paris was arguing about, and the comparison that makes it land is Kenzo, which LVMH had bought three years earlier for around eighty million. WWD ran the story the next morning and mentioned, almost in passing, that the deal had been under discussion since 1994. This page had closed weeks before any of that was public.

There was exactly one precedent for what came next, and it was two months old. Arnault had put John Galliano into Givenchy in July 1995, and Galliano's first couture collection for the house showed on 21 January 1996 to the reviews he had been hired to get. Everything else arrived after this advert went to press. Galliano moved on to Dior that October, McQueen took Givenchy the same month, and 1997 brought Marc Jacobs to Vuitton and Michael Kors to Céline. Arnault was never sentimental about the studios he inherited, having fired Marc Bohan off twenty-nine years at Dior four months into owning LVMH. Céline Vipiana died in 1997, the year her house finally acquired an author.

What's missing from this advert isn't identity. CELINE in capitals across the bottom of the page is a house doing all of its own talking, five cities and a phone number, no interpreter required. What it hasn't got is a signature, and within the year that was the specific thing LVMH went out and bought.

Sources:

This post is timestamped using Blockchain technology. Verify

Larger Than the Evidence

In 2017 I put up four lines about this record, called it haunting and left it there. Accurate, and no use to anyone.

Danny Wolfers, who makes house and techno as Legowelt, wrote the score for Brad Abrahams' short documentary about the Skunk Ape, the bipedal something that supposedly moves through the Florida Everglades. The film runs a shade under ten minutes. The soundtrack is twelve tracks, and Wolfers notes on Bandcamp that most of them never reached the final cut, mostly because there was nowhere to put them, though the sounds would "linger beyond the images of the screen" regardless.

He was right, and the excess is the point. A score far larger than the film it serves is the right shape for a subject far larger than its evidence, which is what I was arguing yesterday about cryptids outliving their own science. Abrahams had a cryptozoologist, an indigenous conservationist and a man who photographs nude models in the swamp, and not ten minutes to hold them.

The cover argues the same way. A red pickup trailing exhaust takes a road through drowned cypress, drawn in crayon like a child's memory of the place instead of the place.

Titles carry it too. "Look For Sign", "Observe Me", "In Their Right Mind", "Forteana", "Swamp Cave" read as field notes from someone who has been out there a long time and found nothing, which is what Abrahams' witnesses mostly describe. The record's own Bandcamp page carries the tags ambient, americana and witch house at once. Witch house, on a Florida swamp documentary, in 2015. Three categories with no business cohering, and nobody involved seems troubled by it.

Sources:

This post is timestamped using Blockchain technology. Verify

Cryptozoology Failed and Cryptids Won

The International Society of Cryptozoology ran from 1982 to 1998, published twelve volumes of an annual research journal, and then dissolved. You can read the scans on the Internet Archive, which is a slightly melancholy way to spend an evening. The society's collapse reads like the end of the story, and Sharon Hill's argument in Contemporary Legend is that it was closer to a release.

The original offer had been zoological. Sightings pointed to unclassified animals, and enough patient fieldwork would eventually produce a body, a specimen number, and a Latin binomial. When the professional structure went, Hill's point is that nobody was left with any standing to police what the word meant. "Cryptid" drifted free of zoology and widened to cover more or less any strange, liminal, arguably-sentient thing you can tell a story about. The sciencey term declined while the useful one spread.

You can see how little the zoology was ever doing by looking at what the field ignored. The historian Mike Dash has noted that few scientists doubt there are thousands of unknown animals still awaiting description, mostly invertebrates, and that cryptozoologists showed essentially no interest in cataloguing any of them, preferring creatures that had defied confirmation for decades. Describing a new beetle takes training and collections access, so this isn't quite a preference. But the incentives never pointed at the achievable version even slightly, and a movement organised around discovery might have been expected to drift that way on its own.

The incentives point at festivals. Point Pleasant, West Virginia, has a Mothman museum and a Mothman festival, and it isn't unusual now: Hill tracks a broad resurgence in town-specific cryptid events, alongside "Cryptidcore" as an online aesthetic people wear as identity. There's real money in it, one 2024 study put Bigfoot-branded products at around $140 million a year, though merchandise totals are soft numbers for a category this fuzzy. The sturdier evidence is that people keep turning up in person. Daniel Loxton, who co-wrote a thoroughly sceptical book on all of this, has said the community-building is worth something on its own terms, which is a generous read from someone with every reason not to be.

The internet-native cryptids are not the same phenomenon, and that's what makes them useful. Slender Man, the Rake, Loab: authored fictions with invented backstory wrapped around them, which the Skeptical Inquirer separates from the older kind by calling fakelore. They make no observational claim at all. Nobody is searching, there's nothing to find, and they circulate anyway, which tells you what the surrounding activity was supplying that the search never had to.

The sociology of Bigfoot encounters closes the loop. To recognise what you saw as Bigfoot you need to know already what Bigfoot is like, how tall, what colour, what noise it makes, so the encounter draws on the folklore and then feeds back into it as testimony. That's a knowledge-making community doing what communities do, which is also why the argument between believers and sceptics never resolves. I've written before about the certainty people hold about things they can't prove. The evidence was never what the certainty rested on.

Sources:

This post is timestamped using Blockchain technology. Verify

Nobody Can Correct Me About 1991

Go looking for proof of a specific afternoon in 1991 and the failure has a texture you don't expect. Gaps, fine, you were braced for gaps. The surprise is that the gap looks identical to the thing never having happened. No partial record, no degraded copy, nothing that survived badly enough to argue with. The query returns nothing, and nothing is exactly what the present returns for events that are simply fictional. Absence of evidence and evidence of absence collapse into the same blank page, silently.

The web's memory has a start date, and it's later than people assume. The Wayback Machine's captures go back to 1996, the year the Internet Archive was founded, and almost everything before that floor is online because somebody later decided to put it there. Which makes it back-fill rather than a record. A 1993 photograph is searchable today because a particular person, at some point after 2004, owned a scanner and had a reason. What survives of that world was therefore selected twice: once by whatever chance preserved the physical object, and again by whoever later felt strongly enough to digitise it. Bands get back-filled. Football clubs get back-filled. Ordinary Tuesdays don't.

None of which means the period went unrecorded. It was documented at the wrong resolution. Newspapers, electoral rolls, planning applications, the local paper's account of a factory closing in March 1992. You can establish the closure to the week and find nothing whatsoever of the people who worked there, which is usually who you were looking for. Individuals appear in that record as a line in a register, present but not described. I've written before about the way information had mass in that period, how knowing something required physical movement. The same physicality governed being known.

The asymmetry that does the damage isn't about volume. The earlier half of a life can't be audited. I can be corrected about 2016 by a timestamp: someone produces a message showing I've misremembered the order of events, and I concede, because the external record outranks me. Very little performs that function for 1991. A payslip can settle a date and an electoral roll can settle an address, so the auditing isn't zero, it's just confined to the handful of facts an institution had a reason to write down. Everything the memory is actually made of, the sequence, the texture, who said what and how it landed, runs unchecked. It drifts the way memory always drifts, and no mechanism anywhere will ever catch the drift. The comparative baseline isn't only unsaved. Most of it was never falsifiable in the first place.

Being fair to the other side of the boundary complicates this rather than dissolving it. The continuous archaeological layer we've been depositing since about 2000 erodes while it forms. Pew Research found that 38% of web pages that existed in 2013 were unreachable a decade later, and that a quarter of all pages from the 2013 to 2023 span have gone. Jason Scott of the Internet Archive put the physics of it well in a piece Adrienne LaFrance wrote for The Atlantic: a piece of paper can burn and you can still get something from it, whereas with a hard drive or a URL, when it's gone there's zero recourse.

The two regimes still fail differently. Modern loss leaves a shape behind. Pew could count the missing 2013 pages because a list of them existed to check against, and a dead link is itself a durable record that something was once there, even when the contents are unrecoverable. That's what makes web archaeology possible at all: Peter Webster reconstructed a late-1990s web sphere of conservative British Christian sites by working outward through hyperlink data in the UK Web Archive, inferring the shape of what existed from traces in the pages that survived. I found that work through a British Library blog post about it, which is now a 404. I checked twice, because it seemed too neat.

For 1991 there isn't even a dead link to fail to follow. Nobody can tell you what proportion of that year is missing at the level of an individual life, because constructing the denominator would require exactly the archive we're saying didn't exist. The people who were there are the index now, and a biological index is unversioned, unaudited, and shrinking by attrition. When you interrogate the present for evidence of that world, the present isn't withholding anything. It was never asked to keep the file.

Sources:

This post is timestamped using Blockchain technology. Verify

A Sample, Not a Faculty

Somebody finally checked the exam paper. Researchers re-annotated 5,700 questions across all 57 subjects of MMLU, the general-knowledge test that anchored nearly every model launch for four years, and found that around 6.49% of it contains errors: wrong answer keys, ambiguous phrasing, questions with no correct option at all. In the virology subset, 57% of the questions they examined were flawed. Correcting the mistakes changed the model rankings, which means part of the ordering we had been reading as capability was models agreeing with a marker who was wrong. Alex Williams collected that study and several like it for Communications of the ACM last week, and a smaller detail in his piece is worse. BIG-bench, built by hundreds of researchers, shipped with a canary string: a unique token embedded in the dataset so anyone training a model could filter the benchmark out, and anyone auditing one could check whether they had. When OpenAI ran contamination checks for the GPT-4 report, BIG-bench had been swallowed by the crawler anyway. The model can produce the canary on request.

Nobody decided to cheat there. The pipeline did it by default, because a held-out test set that has sat on GitHub for three years is not held out in any sense that matters. Instrument noise is a third failure of the same sort: the Leaderboard Illusion authors submitted two identical checkpoints of one model to Chatbot Arena under different names and the scores landed 17 points apart, about the size of gap that gets written up as a generational leap. None of this is subtle, and all of it is fixable in principle. Rotate the questions, proofread the keys, publish confidence intervals, stop reporting single runs.

The objection that doesn't dissolve under better hygiene was made by Raji, Bender, Paullada, Denton and Hanna at NeurIPS in 2021, and I think it's correct and mostly ignored. Treating any benchmark as a measure of general ability is a category error, not a calibration problem. A benchmark is a specific, finite, contextual set of tasks. You can make it bigger, cleaner and fresher, and it will still be specific, finite and contextual. No amount of engineering converts a sample into a faculty. So a cleaned-up leaderboard buys you a more honest number about a narrower thing, and the narrowing is the whole content of the result.

Which is why the measurement work I find worth reading isn't the work that claims to have built a better exam. It's the work that says out loud what smaller quantity it is actually reporting.

François Chollet's version, running since 2019, is that intelligence is not a stock of solved problems but the efficiency with which you acquire skill at problems you've never seen. ARC-AGI is built on the distinction: it tests fluid intelligence rather than crystallized, and restricts itself to a small set of Core Knowledge priors so a system can't win by having read more than the person it's compared against. A model that arrives holding task-specific knowledge the human lacks is scoring the cleverness of whoever encoded it. ARC-AGI-3, released in March, pushes the idea about as far as it goes. Agents are dropped into turn-based environments with no instructions, no stated goal and no reward signal, and have to work out what the game is before they can play it. Scoring compares the actions an agent burns against a human baseline rather than counting right answers. Humans solve 100% of the environments; frontier systems, as of March, score below 1%. The report states the scope plainly, fluid adaptive efficiency on novel tasks and nothing else, and maintaining it is manual work: each version has been rebuilt to resist the optimisation that ate the last one, so what ARC-AGI offers is a gap its authors keep re-opening by hand.

METR changes the unit rather than the questions. Instead of asking what fraction of a fixed set a model gets right, it asks how long a job can be before the model stops finishing it. The 50% time horizon is the human-expert completion time at which an agent succeeds half the time, fitted across a couple of hundred software tasks with real people timed on the same work. January's update grew the suite from 170 tasks to 228 and moved it to new infrastructure. Measured over the full history the horizon doubles roughly every six and a half months; measured since 2023, about every four; since 2024, under three. The doubling time is itself halving, and that is the finding, not the third significant figure METR attaches to each estimate. I like this number better than any accuracy percentage, partly because hours of human work is a unit a non-specialist can hold, and partly because it fails visibly: a saturating suite runs out of long tasks, a shortage you can see in the task list rather than a ceiling hidden in a percentage.

OpenAI's GDPval asks a third thing, whether the model can produce the actual deliverable. It draws 1,320 tasks from 44 occupations across nine sectors, based on real work products, and has experienced professionals from the matching occupation blind-compare model output against human output. It is explicitly positioned against the exam format, which is a fair criticism arriving from a company that spent years publishing exam scores. Its limit is economic rather than conceptual: expert human grading costs money per item forever, so the property that makes it credible is the property that stops it scaling, and the lab funding it is a lab it evaluates. We've already seen how carefully a scoreboard can be arranged when the same party sets the test and reports the result.

These are not competing answers to one question, and reading them as a leaderboard of leaderboards is the mistake. A system can extend its METR horizon by sustaining longer software tasks while staying useless on GDPval's deliverables, and ARC-AGI has nothing to say about either. Adaptation efficiency, autonomous task duration and occupational output quality are three quantities, not three estimates of one. The argument about what the milestone even is persists partly because people keep expecting one of them to settle it.

What I do when I'm choosing a model for real work is duller than any of this. I keep a small set of tasks drawn from work I actually have: verify nine URLs and report honestly which ones are dead, read a thousand-word draft and find the paragraph that sags, take a photograph and say where the faces are. They never get published, so they can't be trained on, and they measure the only thing I need to know, which is whether this model does my job. Public scores are close to useless for the first and silent about the second. That's a purchasing procedure rather than a theory of intelligence, and I've stopped waiting for anyone to hand me the second one.

Sources:

This post is timestamped using Blockchain technology. Verify

Gloves Worn Indoors

Blue, and then more blue, and then the gloves. Powder-blue leather with a green inlay running up the back of each hand, worn indoors, held up at the throat, in a studio where there was nothing for a hand to need protecting from. They match the jacket exactly. The green in them answers the green in the scarf, the scarf answers the gold in the earrings, and the earrings come back round to blue in their cabochon stones. Five colours are working at once, blue, white, gold, orange and green, which ought to be chaos and isn't, because each one has been given somewhere else to be.

January 1993 is a peculiar moment for a photograph like this to be sitting in American Vogue. The appetite for excess is all still here, gold, silk, saturated colour, height in the hair, but the staging has been stripped to almost nothing: a plain near-white sweep, no set, no props, no gradient behind her head. Put the same clothes in 1987 and you'd expect a room, or at least a lit backdrop doing some work. That's one frame and one photographer's choice rather than proof of a movement, so take it as a reading. Still, the ornament stays and the environment goes, which is the direction the whole decade was travelling.

Helena Barquilla's calendar sits on the same line. She is Spanish, and her spring 1993 season ran through Lacroix, Byblos and Hervé Léger alongside the Madrid houses, among them Purificación García, whose best work of that period went almost entirely unwatched. By autumn she was walking Mugler and Krizia and doing couture for Balmain and Dior. She also walked Claude Montana's own-label spring show that year, by which point Lanvin had already replaced him, two Golden Thimbles notwithstanding. That's the constructed, declarative end of the trade, and within a couple of years a good deal of it would read as period rather than present tense.

The makeup is doing the same job as the tailoring. Count the decisions in the eyes alone: a dark brow taken to a point, black liner along the upper lash line, lashes loaded, a shaded socket, a warm sculpted cheek and a terracotta-brown lip drawn rather than smudged. The skin underneath isn't chasing the glassy, half-wet finish the contemporary version of this face would insist on. It's smooth and matte and expensive, and the objective isn't natural beauty at all, it's idealised sophistication, a different product entirely. The hair settles it. Swept back off the face, controlled, with real height at the crown, it supplies the authority the shoulders and the earrings then confirm.

So the result reads less like a woman who happened to be photographed and more like a manufactured image of cosmopolitan adulthood. I mean manufactured as a compliment. Contemporary styling spends most of its effort concealing the fact that styling occurred: the undone hair that took ninety minutes, the no-makeup makeup, the borrowed-from-a-boyfriend shirt fitted to the millimetre. This photograph does the reverse. It announces that a person has been dressed, coiffed, made up, lit and shot, and the announcement is most of the pleasure.

Which is why images like this feel haunting and not merely dated. Old clothes aren't haunting. The thing that has actually gone is the assumption sitting underneath the collar and the gloves, that an adult woman should look completed, and that looking completed was a reasonable public ambition rather than a slightly embarrassing one. Fashion swapped it for looking unbothered. You can date the swap roughly and everybody does, but the more useful measure is that nothing in this frame is apologising for the effort.

The gloves are where I'd start if I had to defend that. Both hands are up at the collar, wrists bent, fingers curled rather than gripping, fingertips resting on the white shirt without pulling at it. Nobody adjusts a collar that way. It's the gesture of someone arriving somewhere or about to leave, borrowed from film rather than from fashion, and it commits the picture to a narrative it never explains. She is not wearing gloves because her hands are cold. She is wearing them because the woman in this photograph is the kind of woman who has gloves.

Sources:

This post is timestamped using Blockchain technology. Verify

Two Person-Centuries

Writing in the Artificial Intelligence Review in 1989, Jordan Pollack got the big thing right and the important thing wrong. He said the connectionist revival would keep growing, and it did, beyond anything he set out. What he expected it to grow into was a field he called connectionist fractal semantics, in which the distributed pattern would become a new kind of symbol, one with internal structure you could reason about, capable of pointing back to the larger structure it came from. He had a name ready for the object: a supersymbol. Nothing of that sort was ever built. The prediction failed in the strangest possible direction, by being fulfilled in volume and refused in kind.

Between 1988 and 1992 artificial intelligence occupied a strange interval between promise and afterlife. One system was fading and the other returning, and for a few years both were audible at once. Symbolic AI still carried the authority of the old dream: intelligence made explicit, filed, named, formalised, replayed. A bureaucracy of mind, in which every concept has its proper label and every conclusion can be traced back along a chain of reason to where it came from. The machine would not learn the world so much as be instructed in it, and for a while this seemed less like one option among several than like the only description that could possibly be true.

By the end of the decade the dream had gone hollow. Expert systems, sold as the industrial future of the field, turned out to be brittle contraptions, impressive inside narrow corridors of competence and expensive to keep upright outside them. They worked when the world behaved like the system's map of the world, and ordinary reality refused that containment, going on producing exceptions, ambiguities, tacit meanings and half-known contexts that no rule base absorbed.

Symbolic AI didn't disappear, though. It lingered the way an institution lingers after its purpose has become uncertain, and the great knowledge-engineering projects of the period have a melancholy grandeur about them. If machines lacked common sense, then common sense could be entered by hand, one assertion at a time, as though the entire background of human life were a document awaiting transcription. Douglas Lenat began Cyc in July 1984 at MCC in Austin on roughly that premise. By the end of its first six years the project had entered over a million assertions, and Lenat's own estimate of what remained was about two person-centuries of further work to reach the hundred million he thought necessary before the system could begin learning usefully on its own. Two centuries of clerical labour, budgeted, to arrive at the starting line. The more the project encoded, the more clearly it showed the abyss underneath knowledge: the residue of habit, embodiment, memory and practical familiarity that people draw on constantly without knowing they are doing it.

Connectionism was coming back at the same time, out of an earlier obscurity. The two Parallel Distributed Processing volumes landed in 1986 and offered a different picture, cognition as pattern emerging across many small adjustments rather than the manipulation of clean symbols. Meaning spread across weights instead of filed in a drawer. The claim that mattered for what followed was not that this worked better. It was that the distributed pattern was supposed to remain readable: microfeatures standing for something, a representation you could open.

The revival had a spectral quality, because none of it was new. Pollack described it plainly as the rebirth of a programme that thrived from the forties through the sixties and was severely retrenched in the seventies. The ideas had been there near the beginning and were pushed aside by the prestige of symbolic reasoning; the key training algorithm had already been written down in 1974 and left to sit. What surfaced in the late eighties was a path not taken, resurfacing precisely as the official future began to decay.

The symbolic order got one more authoritative statement, and it was a good one. Jerry Fodor and Zenon Pylyshyn published their critique in Cognition in 1988, opening with the observation that connectionist models were catching on, that there were conferences and new books nearly every day, and that the fan club included the most unlikely collection of people. Their argument was systematicity: anyone who understands "John loves the girl" understands "the girl loves John", and a classical architecture explains that symmetry for free. What looks like a concession in their paper isn't one. You can reconcile the two, they write, all that's required is that you use your network to implement a Turing machine, which is to say a network only gets systematicity by becoming the thing it claimed to replace.

David Waltz, a year earlier, had doubted you could take a large randomly wired network, show it enough raw sensory input and desired output, and get intelligence out the other end. The learning space for vision and audio was astronomically large, he said, and learning to perceive by feedback seemed cognitively and technically unrealistic. On the narrow question of whether the method scales, he was wrong, and the last fifteen years are the refutation.

The asymmetry between those two objections is the thing I'd point at. Waltz was answered. Fodor and Pylyshyn were not answered; they were outrun. Nobody demonstrated that distributed representations achieve systematicity without implementing a classical architecture underneath. The models simply got large enough that the question stopped being asked, which is a different outcome from being settled, and it left the philosophical objection intact and unattended somewhere behind the industry.

So the lost future of that interval isn't symbolic AI. Symbolic AI failed legibly, and a legible failure can at least be mourned on schedule. The loss that goes unmarked is the connectionism imagined in 1989: the readable microfeature, the supersymbol with inspectable internal structure, the representation you could open and reason about. That research line didn't die so much as get overtaken by systems whose representations nobody can read. We have interpretability now as a field, staffed and funded, which is itself the admission. It exists because the thing Pollack expected to come built in has to be excavated after the fact, from the outside, with uncertain results.

Cyc, meanwhile, never stopped. The transcription continued for decades, and by 2017 the knowledge base held something like 24.5 million assertions, roughly a quarter of the way to the line Lenat had named as the beginning. The project outlived its own future, which is a quieter fate than collapse and much harder to put a date on.

Sources:

This post is timestamped using Blockchain technology. Verify

Fifty-Nine People to Keep It Running

By 1989, Digital Equipment Corporation had fifty-nine technical staff assigned to maintaining the infrastructure and rule base behind its internal expert systems, at that point the most widely publicised application of AI anywhere.

The obvious objection is that fifty-nine people was a bargain. XCON, the configurator most of those rules served, was reckoned to be saving DEC around $25 million a year, and against a number like that a headcount is rounding error. Fair enough. The trouble was never the size of the bill, it was the shape of it: recurring, rising with the rule base, and quoted to nobody at the point of purchase. A system sold on the promise of bottling up scarce expertise turned out to need a permanent staff to keep the bottle from going off.

The winters get told as a story about capability. The machines couldn't do what was claimed, so the money left. Thomas Haigh's reading of the record is less tidy: the famous first winter of the 1970s largely didn't happen, and the real slump was the two-decade one following the 1980s bubble. What collapsed in 1987 was a hardware market rather than a technology. Cheap Unix workstations from Sun ran the same software the specialised Lisp machines ran, and the dedicated machines stopped making sense. Plenty of the expert systems kept running for years afterwards on ordinary computers. The field concluded the idea had failed, which was a larger conclusion than the evidence supported.

Nobody can claim they weren't told. Drew McDermott used the word at an AAAI panel in 1984, at the top of the boom, on a bill called The Dark Ages of AI, borrowing it from the nuclear winter argument then going on. He described a deep unease that the expectations being set would end badly, and he was four years early.

So the lesson everybody agrees on is don't overpromise. Ted Senator puts it plainly in his AAAI paper on what the bust should teach this boom: be measured about strengths and limitations even when the excitement is genuine. That's the cheap lesson, though. No one has ever been talked out of a funding round by their own caution.

The expensive lesson is the fifty-nine. The hardware business died in 1987 for reasons of its own, but what stopped companies replacing their expert systems was never the cost of building them. It was the cost of keeping them correct while the world they described moved underneath. That problem hasn't gone away, it has been renamed. Engineers air it constantly, as complaints about evaluation suites going stale and retrieval indexes rotting quietly. Where it doesn't appear is the investment case, which is still written in training runs and inference margins, as though correctness were a fixed cost you pay once.

The cold periods, on whichever count you accept, arrived when the money behind AI was overwhelmingly governmental, including much of what looked from outside like a commercial hardware market. A handful of decisions could switch it off: the Lighthill report in Britain, the Strategic Computing Initiative cancelling new AI spending in 1988. Henry Kautz argues a third winter is unlikely, and he may be right that the floor sits higher now. Commercial money tied to renewals isn't obviously safer, though. It fails differently, as erosion rather than a freeze, and erosion has no announcement date anyone can point at afterwards.

Sources:

This post is timestamped using Blockchain technology. Verify