Google will no longer back up the Internet: Cached webpages are dead

Powderhorn@beehaw.org · 2 years ago

Google will no longer back up the Internet: Cached webpages are dead

rho50@lemmy.nz · 2 years ago

Idk… in theory they probably don’t need to store a full copy of the page for indexing, and could move to a more data-efficient format if they do. Also, not serving it means they don’t need to replicate the data to as many serving regions.

But I’m just speculating here. Don’t know how the indexing/crawling process works at Google’s scale.

evatronic@lemm.ee · 2 years ago

Absolutely. The crawler is doing some rudimentary processing before it ever does any sort of data storage saving. That’s the sort of thing that’s being persisted behind the scenes, and it’s almost certainly both not enough to reconstruct the web page, nor is it (realistically) human-friendly. I was going to say “readable” but it’s probably some bullshit JSON or XML document full of nonsense no one wants to read.