Future Product
Issue № 001 · 21/08/2026
Discovery

AI Is Eating the Web's Memory. Your Discovery Stack Is Next.

If the open web becomes a write-only feed for model scrapers, product managers own what gets remembered, and what doesn't.

AI-written, human-edited, never fabricated. How this is made

an old library card catalog drawer with handwritten index cards, one card being pulled and retyped into a glowing terminal screen, several other cards fading into blank paper, soft overhead lamp light

The shared archive is thinning, and product teams are about to inherit the bill. Last week, The Walrus walked through a problem product managers usually treat as background noise: as AI systems scrape and rephrase the open web, fewer humans click through to the originals, fewer pages link to one another, and the connective tissue that made the internet a usable reference library quietly dissolves.

You can already feel it in a quarterly review. A launch post gets summarized, the summary is scraped, the scrape is summarized again, and somewhere in the third pass the date, the author, and the source quietly drop out. The result reads authoritative and cites nothing. That is not a search problem anymore. It's a discovery problem, and you're shipping the products that sit on top of it.

If the open web becomes a write-only feed for model scrapers, what gets remembered?

What actually disappears

It's tempting to frame this as a traffic story, the open web losing pageviews to chat answers, and stop there. The more dangerous loss is structural. Search engines used to do two jobs in one motion: answer your question and act as a ledger of who said what, when, and where. That ledger function is what made the web self-correcting, because a weird claim could be chased back to a primary source in two clicks. When AI systems consume pages without users landing on them, the ledger doesn't break, it just stops being updated.

The Walrus makes a sharper version of this point: the page is no longer the destination, it's the feedstock, and once feedstock stops being inspected, it stops being maintained. The original publishers go quiet. The corrections never come. The next model trains on a slightly worse version of the record than the one before. Nobody notices because nobody is reading the underlying pages anymore. Eight months is a long time to not be noticed.

What this means for your roadmap

This is the part where I'd normally hedge, because the connection between a cultural essay in The Walrus and a product roadmap isn't obvious. But here it is, and it's concrete. If AI is becoming the primary interface to information, then discovery features inside your product need to do explicitly what the open web used to do implicitly: keep the citation, keep the timestamp, keep the link to the original. Not as a nice-to-have. As a requirement.

Practically, that means three things on the roadmap you probably already have. First, any in-product answer, summary, or recommendation surface needs a visible source list, and it needs to be one click to the canonical URL, not a hover, not a footnote buried in a settings panel. Second, your internal knowledge features need an archival layer: a way to mark what is canonical, what is a snapshot, and when each was retrieved. Without that, your own team's "ask the docs" feature will start confidently citing last quarter's numbers as if they were this quarter's. Third, provenance has to be a data model, not a footnote. If you can't trace a piece of generated content back to the page that produced it, you don't really know what you shipped.

The honest uncertainty

I want to flag what I'm not sure about, because pretending to be certain here would be a disservice. We don't yet have public measurement of how fast citation chains decay when a category shifts from search to chat, and the labs aren't publishing it. It's possible that internal traffic patterns stabilize at some new equilibrium where a smaller group of power users keeps the ledger alive. It's also possible that the decay is faster than the cultural narrative suggests, because model providers have a strong incentive to compress sources rather than surface them. Both can be true. We won't know for a year.

What we do know is that the products on the other side of this shift are being designed right now, and provenance is one of those features that's painful to bolt on after launch. The teams that treat the citation layer as part of the core data model will ship cheaper, faster, and with fewer compliance surprises in eighteen months. The teams that treat it as a "we'll add footnotes later" ticket will find that later arrives as a security incident or a trust collapse. There's a version of this story where the open web remembers itself, because someone built a product that insisted on showing its work. Build that product.

Sources