Rendered at 06:30:22 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
MetaWhirledPeas 8 hours ago [-]
Here's a contrarian timeline: maybe the old web will return? My reasoning: when the "old web" was great, most people thought the internet was for nerds. Sure even casual users would forward funny emails to their friends, but when it came time for news most people would read the paper or watch the nightly broadcast. When they paid for their Big Mac they'd lay down cash. There were plenty of people using the internet, but again, they were nerds or nerd-adjacent. So, maybe the LLM-ification of everything could end up being like a huge filter, where all the non-nerds no longer see the point of anything and move on to gardens with even higher walls. And then what you have left will be a small subset of people who know how to reach their desired corners of the internet, just like before.
monk_grilla 6 hours ago [-]
We're participating in it right now. HN is a living example of a simple website facilitating discussion between humans with a shared interest. That's why I come here more than any other news/social site now.
tavavex 5 hours ago [-]
I didn't live long enough to see the entire lifetime of HN, but it's definitely part of the old web in my book. Its functionality is the same it's always been, it's rock solid, lightweight and simple. No personalization or engagement maximizing, no modern-style ads. Being spun off from something as big as Y Combinator meant they can run this site almost indefinitely.
ErigmolCt 1 hours ago [-]
HN is a good example of how little a website actually needs once you remove the requirement to constantly squeeze more engagement out of it. Though to be fair, its audience is unusually self-selecting :)
sodapopcan 4 hours ago [-]
The only thing not "old" about it is upvoting, which is one of those YMMV features.
marcosdumay 3 hours ago [-]
It has the entire web 2.0 neophillia. Discussions stay for a day or so in focus, then they are gone.
This one is almost gone already.
samudrijan 1 hours ago [-]
Were you even here?
CoastalCoder 4 hours ago [-]
> I didn't live long enough to see...
Whoa, careful with your verb tenses! You just wrote a ghost story :)
tavavex 2 hours ago [-]
Have seen. Too late to edit it now...
apsurd 11 minutes ago [-]
havent lived. “I didnt live long enough” means your life ended! past tense. hasn’t means its still going “have not (yet)”. I don’t know the formal rules behind it all though.
its interesting because it sounds more correct to use have not yet because your life is still going on but its going on in the forward direction, not backwards in time to overlap an earlier HN state. So you could say its not right, but the sentence sounds/feels more right to preserve that your life is still going on.
We like that you are still here with us.
zem 5 hours ago [-]
what would be amazing would be a resurgence of the old style forums, but where users could choose to see a flat or tree view of a discussion (and even customize those views, e.g. the tree could be reddit style without subject lines, or slashdot style with them)
amazingamazing 5 hours ago [-]
A simple website controlled by VCs, lol
Xorakios 5 hours ago [-]
agreed
sieabahlpark 6 hours ago [-]
[dead]
doright 4 hours ago [-]
I don't believe the old web was filled with people pining about how the current web is awful and them participating in recreating the web they inhabit is better. People on the old web just posted things, because before that there was no non-old web to compare to. But that in itself is what makes the old web so appealing to people wanting it to return. Back then there was no other web necessary to escape from.
So when people construct versions of the old web today, they make a conscious decision to do something most people would not, and by doing so it sometimes comes off as indicating what they think of the other design decisions they chose not to make. It's not a bad thing, but it's not quite the same as it was anymore. That is why some of the spark is gone, in my eyes. It's kind of like counterculture being mistaken as the opposite of culture.
It is even more pronounced when people talk about "the old web" on old web websites all day.
kalleboo 55 minutes ago [-]
> I don't believe the old web was filled with people pining about how the current web is awful and them participating in recreating the web they inhabit is better
TBH I do recall a lot of gloating over how AOL was full of clueless lusers who were inferior to the Internet
goodmythical 7 hours ago [-]
I haven't checked it out personally, but this is my impression of all of the webs that aren't the world wide web. Like the meshtastics, the new freenet, some web 3.0
It being harder to join has the effect of their being a smaller more engaged user base. Less likely that there will be pressure from corporate interests, more likely that every user has a vested interest in both production and consumption.
edit: And of course, you'll see the same in dedicated niche forms and spaces like HAM radio. I imagine the folks TXing EME have a particular flavor that is well maintained.
6 hours ago [-]
ErigmolCt 1 hours ago [-]
I think the audience could absolutely shrink back to something resembling the old web. I'm less sure the old web itself comes back
zem 5 hours ago [-]
the article already lists one positive - the LLM assistance made the engineering effort to keep their site up and running feasible again. insofar as spam killed the remnants of old web, that is a promising sign.
cookiengineer 3 hours ago [-]
Your misjudgement comes from not seeing the economic incentive here. The web as a surveillance mechanism is just way too juicy to let go. As of now, pretty much all economies depend on it, just from an economic numbers calculation.
Mind you, 200 years later, we still live in cities built for cars, not people. Rockefeller, Shell, BP and the like are still not dead, and still use everything they got to not go extinct while corrupting every politician that will accept their money. Just for a couple more years at a time, and a couple more, and a couple more...while destroying our planet like a cancer.
In 200 years we'll still have the web built for surveillance capitalism, not people.
There is no old web to go back to. Meanwhile state surveillance depends on its existence. Amazon, Apple, Google, Meta, Microsoft, three letter agencies, Palantir, and now the AI companies depend on it, and they're selling to nation states at a time. Again, it's just too juicy not to.
Back then we called them trusts, but now those big tech companies are seemingly worth more than actual countries. That's too much power without legislation, as they can effectively abandon the rule of law. Remember what happened with the Enron scandal? That was about billions, not trillions.
If we want a human web again, we need to build a new one, with privacy and secrecy first and not as an afterthought. No CA, and with mutual cryptography that acts like both a ledger and as a privacy shield for the users. No cookies, no tracking, no degenerate google controlling it all through owning the source code.
Razengan 7 hours ago [-]
Allow me to switch to asshole mode for a minute:
"Things used to be so good" is mostly rose-tinted glasses: What exactly WAS the "old web" anyway?
100s of half-baked sites hosted on Geocities, Yahoo, about pointless stuff? Epilepsy simulator websites splattered with gif-vomit? Everyone and their grandmother asking you to install their 10 toolbars?
Honest, serious question: What actually was of objective substance on the old internet that's nowhere to be found now?
You can find random pointless stuff now too, just that except Geocities/Yahoo it's Intsagram/TikTok/Twitter etc.
If you mean self-hosted websites, they're still here, and nothing's stopping you from creating your own.
If you loved all the Flash toons on Newgrounds etc there's unironically a lot more shorts and animations on YouTube or even Vimeo now, if you but search for them (the old gods like Weebl, David Firth are still there, and I suggest Sechi, DoodletmeGo, Nondescript Video Club, and からめる to start with then let the algorithm soak up the weirdness)
Just as the "internet" supplanted newspapers, magazines, and broadcast television for many people,
and how the newspapers replaced the town criers before them,
why shouldn't the "internet" be supplanted by a more accessible medium?
There's so much crap on the "old" and "middle age" internet: ads, scams, anal mods, trolls, bandwagons, vote wars, low-effort content and other repetitive noise that could be bypassed by just asking AI about something and make better use of your limited mortal lifetime, like watching that anime you've been putting off for years.
greazy 6 hours ago [-]
For me that's not the old web.
The old web was local forum created by gamers for gamers who hosted real life events. You could post asking for super specific advice to your locale. It was eaten by reddit.
There were cool IRC servers for all sorts of niche projects and topics. This still exists of course but a lot of it was eaten by Discord or Slack.
While I never participated, sites like LiveJournal provided both a unique outlet and view into other people's lives. I think a lot of social media like Instagram and Facebook eat this up.
I'm sure everyone has their own versions. And this is the old web that was gobbled up by mega sites which aim to commericalise the crap out of everything.
Btw I consider HN part of the old web.
sodapopcan 6 hours ago [-]
I loved Diaryland since it was fully customizable. It was my intro into templating! It was very crude, but it worked.
Razengan 6 hours ago [-]
Oh my gosh I loved Diaryland, and how you could customize your shit with some simple HTML, but I think they went through enshittification while they were still around anyway
But..what says something like Diaryland can't be brought back now?
sodapopcan 6 hours ago [-]
Did they? I think it was LiveJournal that enshittified. Diaryland is actually still around. I just checked it out and it now "Diaryland v2!" I actually last checked it out in the past year and it still looked exactly like it did back then!
rockskon 5 hours ago [-]
Newer websites are filtered through what is acceptable from a corporate lens.
The old web culture can't exist because advertisers won't let it exist on modern platforms.
Additionally the spread of information is significantly less organic because of relentless pushes to use algorithmic feeds along with far worse distractions from content than there used to be.
Razengan 4 hours ago [-]
> The old web culture can't exist because advertisers won't let it exist on modern platforms.
Again, what's preventing you from buying a domain name and putting up your own website? bypass the advertisers?
"Not enough people will see it", you say? So you want the reach provided by "modern platforms" but you don't want their rules?
ColdStream 6 hours ago [-]
I think it is a case of, it is easier to find those niches nowadays but when you do it is a more shallow and controlled version of said thing.
On the old web, you would just stumble onto something deep but the barrier to get to it was fairly high. It could also just be total nonsense but at least you could see the passion behind it.
I agree with your general take. People talk about the old web as 'the good old days' but it really wasn't that great. I do have some fond memories of the mid 2000's when it was all starting to mature but also before corporate capture took hold. Even then, you would just end up stumbling around blindly most of the time.
AuthAuth 6 hours ago [-]
Niches nowdays are far deeper and more vibrant than they were 20 years ago. Instead of 1 guys website or 1 fourm with 30 users you now have many discord servers with 1000s of users talking in real time, youtube channels putting out videos and making money, reddit/lemmy aggregating content, the community interacting in real time and building what it needs.
inigyou 6 hours ago [-]
And they have to self censor so they don't get banned off those platforms.
There are so many piracy discords where you can't ask how to pirate something because that's against discord rules. I mean really? Why are you using it then?
AuthAuth 2 hours ago [-]
Most niches dont have to self censor. Even the ones that do still have no trouble communicating their message. I follow a lot of military channels and they post uncensored versions of their videos off yt but there is little reason to go because seeing the gore is not why im watching the video. I still get the entire value of the video from youtube even though you might argue its self censoring.
and don't forget Big Billy's "no more tears" guarantee on all anal mods purchased from the official BBAM shop!
sodapopcan 6 hours ago [-]
> What exactly WAS the "old web" anyway?
Pre-Google, specifically, pre-PageRank, and even more specifically, pre-SEO.
Your post is oddly hyper-focused on pointing out the negatives, probably due to the asshole hat you put on. There were lots of good sites with tons of good, free information, and because the main way people searched was using directories, we weren't inundated with 1,000s of results on the same topic, vying for a top-placement. The writing itself was much more laid back because SEO-optimization didn't exist. Also, those "half-baked GeoCities sites" were awesome! Of course there were a lot of junk ones, but again, it was just a bunch of people writing about themselves, their interests, and generally being creative. And again, NO ONE WAS TRYING TO MAKE MONEY, they were just doing it. No "like and subscribe." No gaming an algorithm. I used to love finding guitar tabs on those sites and I had a couple of sites with lots of tabs I did on my own for two different bands.
> There's so much crap on the "old" and "middle age" internet: ads, scams, anal mods, trolls, low-effort content, that could be bypassed by just asking AI about something and make better use of your limited mortal lifetime.
I feel like this was GP's point, no? Because people are just using AI for search now (this is the main reason I use it) that there will be no point in making these spamming sites anymore, which "opens up" the internet a bit for those of us who enjoyed the old way. Yes, it's all still out there, but it's kind of hard to find, and it would be kinda neat if more people discovered making their websites outside of walled gardens.
Razengan 6 hours ago [-]
> Your post is oddly hyper-focused on pointing out the negatives
because people are hyper-focused on the negatives of the "current" web, or acting like the sites they used to love can't be recreated anymore..
> it was just a bunch of people writing about themselves, their interests, and generally being creative.
There still are. on YouTube, TikTok, Instagram, Reddit, 5ch and other Japanese subcommunities.. the fact that they're all "silo'ed" into these less-than-benign corpoctopuses is another matter, but people are still people'ing
sodapopcan 4 hours ago [-]
> because people are hyper-focused on the negatives of the "current" web, or acting like the sites they used to love can't be recreated anymore..
Fair!
> There still are. on YouTube, TikTok, Instagram, Reddit, 5ch and other Japanese subcommunities.. the fact that they're all "silo'ed" into these less-than-benign corpoctopuses is another matter, but people are still people'ing
I know there are, but as per my comment, it's just harder to find and you used to almost not be able to avoid it. I don't think platforms like YouTube and Instagram are great examples due to the monotization... everyone is trying to "make it." But ya, I mean, I do know it's out there. But I think the whole point wasn't that this stuff doesn't still exist, just that maybe it will become the more prevalent thing again. At least that's how I read GP's comment. But who knows.
zem 5 hours ago [-]
it's not "another thing", it means that your community forum lives on reddit (or whoever's) sufferance, and is subject to some corporation's capricious and arbitrary whims.
Ya know a while back I intentionally locked myself out of my HN account so I'd stop egnaging with the drivel. But you have such a delightfully awful take I had to respond. I can't tell if you're trying to gaslight folks, you're too young to have used the "old internet", or if you're just a true believer but here goes.
Honest, serious question: What actually was of objective substance on the old internet that's nowhere to be found now?
Ah yes the time honored tradition of JAQing off. What's missing? Fast load times. Everyone's hiding behind anti-bot crap now. Either because AI DDoS bots are making hosting unbearably expensive or because folks don't want their content being poorly repackaged for a profit. So now you have to go through some nonsense proof-of-life crap that takes a few seconds before you can even start to load the content. Anonymity. Used to be you could browse sites without having to log in. But in an attempt to thwart greedy AI companies all sorts of public sites like Reddit, Github, and Youtube require you to log in before doing much of anything. Accurate search results. I've yet to see Google or DDG return a remotely relevant AI response and actual search results are absolute garbage now. DDG in particular is putting drek like "DeepWiki" at the top of its results.
If you mean self-hosted websites, they're still here, and nothing's stopped you from creating your own.
Except now you have to contend with AI DDoS bots which leaves people to either stop putting up small self-hosted sites or put them behind centralized anti-bot protective services.
If you loved all the Flash toons on Newgrounds etc there's unironically a lot more shorts and animations on YouTube now
Except it's all buried underneath an avalanche of AI slop. If it's there you're going to have to put in a ton of effort to find it. For the most part I've stopped searching for things on Youtube and instead check in on known/trusted creators periodically. Discoverability on YT is terrible these days.
Honestly I loved having domain specific forums that weren't beholden to a unified shareholder friendly content policy. Now they've been supplanted by a Reddit that's trying desperately to monetize content.
why shouldn't the "internet" be supplanted by a more accessible medium?
AI makes things less accessible because it centralizes power and increases the cost of publishing content.
There's so much crap on the "old" and "middle age" internet: ads, scams, anal mods, trolls, low-effort content, that could be bypassed by just asking AI about something and make better use of your limited mortal lifetime.
Anal mods is oddly specific, if you don't like butt stuff you could theoretically just ignore it. But okay. My time is better spent reading and understanding than waiting for some anti-bot crap to finish before I pick through largely irrelevant AI slop masquerading as an accurate summary.
Those scams and that "crap" largely still exist on the "new" internet. Only now you're going to have to prove you're a human to view them. Where things have evolved is that instead of human written SEO drivel you've now got AI generated SEO drivel. It's to the point that I have to add "-deepwiki -ai" to pretty much every DDG search.
The problem is that AI companies are going full scorched earth. AI provides dog shit answers while ruining traditional search results. Or if you'd like a physical example they're physically destroying rare books to come up with poor, largely irrelevant AI drivel for the masses.
morganf 9 hours ago [-]
I'd define it as the web until the time that Facebook truly took off and conquered the hearts and minds of so many (that was the first huge shot in the losing war of the old web). For example, a key part of the old web was what used to be called the "blogosphere", and the blog's height was those final years before FB (in the ascendancy) and the first years of the FB era (descendancy).
adrianwaj 2 hours ago [-]
For me the new web came about when if you wanted a website, you hired a web developer. As soon as it become DIY, without your own domain and didn't involve looking at HTML, the new web was born.
Today, if you own your own domain and fiddle with HTML then basically you're still part of the old web.
So basically, the old web required experts to help people get online, the new web negated them through homogenization.
The real experts from the old web basically became startup founders to monetize the homogenization and automation of their worldviews for the less tech-savvy (or specialized) folks.
People pine for the old web, but it was a churning, friggin' mess too. Search results were good though, but only when Google arrived. The standardization and ease-of-use from the new web has been a massive boon. I think things will change as people yearn for a wild-west, where there are new opportunities. The promise of a better life, seeing that the grass is greener, avoiding pain - those things always create golden ages.
Once AI favours people who can afford to pay for the best results.. or surveillance becomes too burdensome or dangerous, then watch this space.
The spirit of freedom always runs into trouble, when considering the notion of:
"Cut off the heads of others to make oneself appear taller."
Surveillance enables this. Decentralize again, but how?
ISPs - help or hindrance? Why/how can ISP charges be diverted to content creators?
jperras 10 hours ago [-]
I would posit that "old web" could be defined as the period before Google Search became public (<~1997).
But maybe that's more a measure of my own age and perceptions rather than an accurate representation of the various eras of the internet/web...
binarymax 7 hours ago [-]
That’s the Cretaceous web. I miss it fondly.
Xorakios 5 hours ago [-]
Oh good golly miss molly. I had my first email address at an .edu in 1980, first website in my own real name since 1993. the 2000's are just new kids on the block
ColdStream 5 hours ago [-]
I'm not sure I miss the web as it was but miss the feeling of it all being novel and exciting. You can still get glimpses of it nowadays but it isn't often.
Thanks for the link. In ‘96 I only had text-only/Lynx, so rarely/never got to see what the early versions looked like with a “real” browser.
bradley13 10 hours ago [-]
2009-2014? That's not the old web. Or I'm old. Take your pick.
Gormo 10 hours ago [-]
Yeah, I'd say "old web" is early '90s up through about mid-2000s or so. Geocities, Tripod, Angelfire, lots of standalone web forums, the early blogosphere, no social media, etc.
clickety_clack 9 hours ago [-]
Pre-social media is probably the watermark. That’s what sucked all the content out of the web and into walled-garden platforms.
ColdStream 5 hours ago [-]
As they say, now it is five social media platforms posting copies of each others content.
s0rce 7 hours ago [-]
I think there are a few eras. Pre-graphics, then pre-Google search, then pre-social media/public facebook.
ColdStream 5 hours ago [-]
There was the sub era of Flash based nonsense. Somewhere in the 2002-2005 period.
ErigmolCt 1 hours ago [-]
Yeah, 2009-2014 feels more like the old social web to me
My first grandchild was born in 2000. And the COBOL panic for Y2K is probably partly my fault because 4 characters were too computational expensive when I wrote code that probably still powers your Igloo ice containers and GEICO insurance policies. 80's tech still persists!
mrexroad 9 hours ago [-]
‘09-‘14 is hilariously not the “old web.” With that said, those of us who participated in what I consider the “old web” are likely “old” now, so… shrug?
Two of my favorite sites [0][1] still online—Lurkers Guide to Babylon 5 and ex-astris-scientia—were started in ‘92 and ‘98 respectively.
1992 would be super SUPER early for the web. Looks like it launched in 1994.
mrexroad 8 hours ago [-]
Yeah, I took a shortcut by just grabbing the copyright on page, but Lurker’s Guide to Babylon 5 has a slightly more storied history than being a www site. It started on Usenet before the pilot even aired iirc. I assume the ‘92 date encompasses the early content that evolved into the site it became.
With that said, its domain has changed a few times iirc… so maybe not the best example given the nature of the article.
stillpointlab 10 hours ago [-]
I've just gotten used to this at this point. My first time on the web was somewhere around 1995 I think, although my first time on the Internet was earlier (it was a proxy through some ancient BBS and I'm pretty sure it was using gopher). Even though I was just a kid back then, clearly that makes me old now.
kQq9oHeAz6wLLS 9 hours ago [-]
Speaking of gopher, I'm low-key hoping that becomes the new place for all the non-bot traffic. Gopher felt magical back in the pre-www days.
mattkevan 9 hours ago [-]
I’m old and it’s definitely not the old web. I consider the old web to be when PNGs were sliced with Fireworks and laid out in tables.
2009-2014 is Web 2.0, back when people still thought social media was a good idea.
toast0 6 hours ago [-]
> I consider the old web to be when PNGs were sliced with Fireworks and laid out in tables.
PNGs? We had (mostly static) GIFs and we liked them.
pluc 9 hours ago [-]
Old web was 88x31 buttons, 468x60 banners and tables
b3ing 7 hours ago [-]
2003-2009ish was Web 2.0, part of it was CSS/web standards and AJAX and blogging becoming more mainstream
benjaminl 9 hours ago [-]
For me the old web is when people still had homepages. When those went away, the old web died.
ErigmolCt 1 hours ago [-]
2009 is old web now, and this is how we find out we're old... Yet I think maybe "middle-aged web" is the more accurate term
drooopy 10 hours ago [-]
Right? When I think of the old web I think of my Xena and X-Files fan sites on Geocities circa 1997
darknavi 10 hours ago [-]
I was a few years late to the start (early 90s kid) but I am nostalgic for the vast amount of myfreewebs + dot.tk websites out there.
Hit counters on the front page was mandatory of course.
nkrisc 10 hours ago [-]
It was still a significantly different era than today. Let’s call it Middle Web. Maybe Late Middle Web.
ajsnigrutin 10 hours ago [-]
+1 for this
I'd consider the facebook and mega era to be relatively new, the "old web" for me would be the one without centralization around a few giants, the era of random phpbb forums, private websites with "this site is under construction" banners and internet directories to find stuff.
dd8601fn 10 hours ago [-]
God… I stood up so, so many phpBB instances.
Anon4Now 6 hours ago [-]
When I think of "old web", I always think of "Britney Spears' Guide to
Semiconductor Physics". Thankfully, it still lives:
I think of "Britney's Spheres" which was rather distasteful iirc
Anyway Britney's 1998 debut album predates The Matrix, includes the song Email My Heart and came on "Enhanced CD" with tags and video. How about that
dbetteridge 6 hours ago [-]
This is beautiful Thankyou
z_rho_one 11 hours ago [-]
Remember the good ol' days when we all thought that everything on the web would exist for eternity and over.
ryandrake 8 hours ago [-]
I remember the naive implication/expectation that a URL would be permanent. That once a file is identified by its URL, you could bookmark it and it would always be there. And absolute worst case, if someone really, really, really had to change a file's URL, they would politely return 301 Moved Permanently, and feel very bad about it.
Now people don't give a shit about URLs. Webmasters casually move files around all the time because they feel like it, and if links get broken, who cares, that's the referring site's problem! Their beautiful file hierarchy is more important than the web staying connected!
tesin 7 hours ago [-]
"Webmasters" - now there's a term I haven't heard since 1998!
inigyou 6 hours ago [-]
It's been 40 years. Half the webmasters are dead. You can't keep things alive longer than yourself.
nephihaha 11 hours ago [-]
No, that's just embarrassing stuff. That stays on the web forever.
SAI_Peregrinus 11 hours ago [-]
Yep, it's Murphy's law of online content. Anything you want to reference later will be gone, with no archive copies. Anything you want deleted will be available forever.
johnnyanmac 10 hours ago [-]
Interesting blog post comparing the culture and incentives of CEOs in different countries? Completely gone, tried to search it up multiple times to no avail. It's barely 3 years old.
That random 20 year old video of some middle schoolers doing a flying kick and breaking a vending machine? Yeah, just popped up on my feed yesterday.
11 hours ago [-]
mohamedkoubaa 10 hours ago [-]
Especially since the Internet is literally a _messaging_ protocol.
alex1138 4 hours ago [-]
And then Yahoo shut down Geocities and fought archivers and then shut down Yahoo Groups and fought archivers
11 hours ago [-]
Gormo 10 hours ago [-]
I mean, it does usually, just not always in its original location.
phendrenad2 7 hours ago [-]
It seemed reasonable, because keeping things up on the internet was becoming cheaper. We failed to see that capitalism would even demand payment for such trivial things.
mryall 10 hours ago [-]
Quite ironic that a link shortener which went offline for a decade or so is now posting about other sites not staying online.
The plot features two American men who stumble upon Brigadoon, a mysterious Scottish village that appears for only one day every 100 years; one man soon falls in love with a young woman from Brigadoon. The show's song "Almost Like Being in Love" subsequently became a standard.
ErigmolCt 1 hours ago [-]
The interesting number here isn't 76.7% dead, it's that 23.3% alive is an upper bound. A server answering HTTP doesn't mean the thing somebody shared in 2011 still exists
6c696e7578 10 hours ago [-]
It seems putting anything on the web that allows submit is screaming to get spammed these days. Is there any sort of spam filter that's worth using?
tokai 12 hours ago [-]
Am I getting old? 09-14 is not even close to the old web for me. The old web, to me, was back when people still published physical 'phone' books for websites.
acheron 12 hours ago [-]
Seriously. 2009 is several years after everyone was already saying “web 2.0”! That is nowhere near the “old web”.
rdmuser 11 hours ago [-]
The old web doesn't necessarily mean the oldest web. 12-17 years ago was very much an older fairly different era of the web that's worth analyzing even if it's on the younger side of the old web. I can definitively sympathize with your reaction though, it doesn't feel like that era was that long ago yet.
dasil003 11 hours ago [-]
I vaguely recall those, but they were more for normies trying to get online. For me the old web is what I saw when I logged into my university gopher server and saw the advertisement for something called the World Wide Web which I could browse via lynx. Soon enough I got a PPP connection and then Mosaic/Netscape 1.0. However everything after javascript shipped (let alone CSS) is new new new. I'd almost go as far as saying if it doesn't have a tilde in the URL it's not old web... almost...
dd8601fn 9 hours ago [-]
I’m vaguely remembering tilde username for our public html folders. Was that old Apache behavior?
And now I wonder if Apache is even still common. I spent so much time fiddling with apache configs.
dyauspitr 9 hours ago [-]
That’s the prehistoric web or maybe that’s BBSs…
dasil003 6 hours ago [-]
Correct. Don't get me started on Fidonet, as a young computer nerd in the 80s before even dial-up internet was available, that was a portal to the global internet that my adolescent shit-posting self definitely didn't deserve.
inigyou 6 hours ago [-]
Someone called that the Cretaceous web. The old web was the era before social media sucked everything up (so 2009 wasn't it).
hmhrex 11 hours ago [-]
Purevolume mention made me sad. I miss that community.
7 hours ago [-]
levocardia 11 hours ago [-]
100% AI generated text. The irony.
CqtGLRGcukpy 11 hours ago [-]
Proof? Using an AI checker isn't accurate.
randomblock1 11 hours ago [-]
The ENTIRE thing is AI generated. I'm not talking about the article. I'm talking about the entire website, the entire "product". https://0.mk/blog/zero-humans
Also it's super easy to tell by looking at it, way too many LLM-isms. No need for a AI checker tool.
Levitating 11 hours ago [-]
> The name was registered in 2009 because it was the shortest URL possible: a zero, a dot, two letters. Seventeen years later the zero means something else.
Did the zero ever mean anything? It's still a 3 character domain regardless if the first character is a zero.
shevy-java 12 hours ago [-]
Webpages dying is probably one of the biggest design flaws of the original web.
I am not saying old content needs to be preserved forever, but so much content has factually been lost over time. Old logs from text-based MUDs for instance, even for MUDs that still exist today.
efskap 12 hours ago [-]
It's the great irony of digital media. Copying data is accurate to the bit and is preserved "as-is", but in practice, it requires someone to maintain servers, to care about it. To separate out what is worth preserving.
We hot-linked to all those image hosts because we couldn't imagine them disappearing.
Archive.org had incredible foresight and if it didn't already exist, I'd call such a project a pipe dream.
sumtechguy 11 hours ago [-]
Even Archive.org is rather limited in what it keeps. I know of a very large site that recently disappeared. archive only has part of the web html part of the site. Everything else is either gone or non accessible.
Levitating 11 hours ago [-]
That's not my experience
-0_0- 7 hours ago [-]
There needs to be some kind of "murphy's law" for this style of comment. "For any comment on the internet where someone points out an issue they've encountered with technology, there will inevitably be a reply from someone else sharing that they haven't personally experienced it."
OuterVale 2 hours ago [-]
I think you get to coin it.
sadlyuramoron 8 hours ago [-]
[dead]
cortesoft 12 hours ago [-]
> We hot-linked to all those image hosts because we couldn't imagine them disappearing.
No, we hot-linked all those image hosts because we didn't want to pay to host it ourselves.
EvanAnderson 11 hours ago [-]
...and I had fun replacing images people directly linked from my server with less-- ahem-- savory images.
I enjoyed the emails I got from a couple people who were adamant I "hacked" their site because their "web developer" linked to images on my server... images that now said stuff like "I'm a loser bandwidth thief!", etc. (I never did use really nasty "shock" images-- just taunting stuff.)
marginalia_nu 10 hours ago [-]
> It's the great irony of digital media. Copying data is accurate to the bit and is preserved "as-is", but in practice, it requires someone to maintain servers, to care about it. To separate out what is worth preserving.
This is a bit idealized. In practice copying data is not quite accurate (especially in bulk) and bit-rot is a very real phenomenon, both in flight and in storage.
You sometimes encounter it when dealing with files from the early '00s, it's very common to discover a few of them are corrupt, even if they've only ever been copied between harddrives.
tekne 10 hours ago [-]
Content-addressed storage and error correcting codes mean that one can make bitrot astronomically unlikely with honestly minimal infra investment.
It's copyright that causes anything to disappear from the web IMO -- torrents never die.
EDIT: I am aware that unseeded torrents do in fact die. But it really doesn't take much to seed a whole hard drive's worth of rarely requested data -- this also detects bitrot and so corrects errors automatically if you're not the only copy.
If you are, there's ECC, as well as making another copy.
marginalia_nu 9 hours ago [-]
There are mitigations in both software and hardware, but most consumer machines, by default, do almost none of that. No ECC RAM, no error correction in the filesystem.
There are many TinyMUD logs that were posted on Usenet, still to be found on Google Groups.
However, logging was controversial amongst mudders. It was almost always rude to log a private conversation without knowledge or consent; it was tacky to indiscriminately log while everyone was in the "hangout room" or Rec Room, as it were, and it was also bad form to post logs to Usenet or share them without redacting player names and other things.
But logging was built-in to most clients, and it was possible for server administrators to log (and hypothetically any malware-in-the-middle could log the cleartext, unencrypted TinyMUD TCP streams.) And many nefarious deeds by nasty players were exposed to the light when their logs were posted.
ChadNauseam 12 hours ago [-]
The technology is still in its infancy unfortunately, so there's no way the web could have been based on it, but I think content-addressing is the long-term play. If I click a link, there are some cases where I want the server to respond with a fresh response just for me (e.g. a website showing the weather). But often I just want whatever content was linked to (e.g. a webpage explaining a math content). In the latter case, it would be nice if the link had a hash of the content in it, and 3rd parties could host copies to keep the link working even if the original operator stopped existing.
Gormo 10 hours ago [-]
That's pretty much how IPFS works.
inigyou 6 hours ago [-]
Did they ever solve the problem where requesting a file took ten minutes
rcxdude 12 hours ago [-]
It's pretty difficult to avoid without very significant tradeoffs, though. The closest is content-addressable peer-to-peer networks, but these still rely on someone keeping the information around, and they struggle to scale anywhere near as much.
MetaWhirledPeas 8 hours ago [-]
> Webpages dying is probably one of the biggest design flaws of the original web.
Let me introduce you to the alternatives: print media, film, stone engravings. That stuff tends to get burned and shattered and it takes FOREVER to make copies.
I'm being cheeky but I don't know what design change you could possible make to the web to make webpages not die.
Gormo 10 hours ago [-]
> Webpages dying is probably one of the biggest design flaws of the original web.
I'd say it's one of the biggest design flaws of the current web, what with more and more content hidden behind paywalls, increasingly restricted WAFs, and rendered client-side via convoluted JavaScript.
Archiving and mirroring of old-style websites, delivered as static HTML, is simple and straightforward. 20 years from now, most web content from ~1996 to ~2015 will still be accessible, but much of today's web content probably won't.
twotwigs 8 hours ago [-]
Fun experiment:
Have an LLM “guess” random URLs seeded with words from a dictionary, iterating over each word and guessing a URL.
It guesses a lot of correct URLs. This is one method of “URL hunting” that doesn’t involve a 3rd party list or index.
Then just scan those pages for other URLs, visit them, and add a tally every time you come across a URL (for page rank).
Then search anything, see what the results are. You have invented a dark web search engine.
inigyou 6 hours ago [-]
> dark web
Guessing random words and .onion at the end probably won't get you anywhere, although onion is also quite a random word...
orangenectar 3 hours ago [-]
[dead]
twotwigs 4 hours ago [-]
Where did you get "onion"? Reading is not your strength, I take it
The deep web,[1] also known as the invisible web[2] or the hidden web,[3] is the parts of the World Wide Web whose contents are not indexed by standard web search-engine programs.[4] This is in contrast to the surface web, which is accessible to anyone using the Internet.[5] Computer scientist Michael K. Bergman is credited with inventing the term in 2001 as a search-indexing term.[6]
twotwigs 3 hours ago [-]
That's the one
12 hours ago [-]
hmartin 12 hours ago [-]
Site got hugged? Is there a torrent?
Lord_Zero 11 hours ago [-]
The blog mentions "0.mk's revenue did not cover hosting" but then goes on to implement expensive AI integration. Not counting cost for tokens to do the development.
Also:
> Reply to any 0.mk email and the message lands in a feedback queue the AI reads, triages, and acts on
Is this dangerous? What about jailbreaking AIs and having it delete everyone's account?
Gormo 10 hours ago [-]
> The blog mentions "0.mk's revenue did not cover hosting" but then goes on to implement expensive AI integration.
The article doesn't seem to describe the cost of the AI solution. It does imply that it is lower than the cost of maintaining and supporting their service manually.
exitnode 12 hours ago [-]
Wow, that is a great domain!
SwellJoe 11 hours ago [-]
"The old web" is 1993 to 2007. It's all been downhill ever since.
I found an old database backup of 0.mk on a disk I had kept.
0.mk started in 2009 as a passion project built by three of us. We worked on it for a few hours each week around our regular jobs. We eventually closed it in 2014 because the revenue (hint: no revenue) could not cover hosting, development, and the constant work of fighting spam and reviewing abuse.
The recovered historical corpus contains 657,607 links. For this analysis, we followed every one of them.
Of the 655,178 links with safe, crawlable targets, 76.7% no longer returned a loading page. After removing repeated destinations, 78.7% of the 492,620 distinct crawlable URLs still did not load. So duplicate links are not creating the result.
I use “did not load” rather than “gone” deliberately. Some URLs returned 403 or 429 and may have blocked the crawler. Pages that returned 2xx or 3xx count as loading even when they now lead to parked domains, login walls, or removed-content notices.
There is one large distortion in the yearly data. A single account created 83,398 URLs pointing to one hostname in 2011. At URL level, 92.5% of that year did not load. Count each hostname once and the result becomes 61.7%, almost identical to 2010 and 2012.
A few things I did not expect:
- 835 restored links point at Facebook’s old photo CDN. None loaded.
- The first link ever shortened was a CSS stylesheet on a WordPress blog.
- Someone shortened localhost on the second day.
- The longest stored URL is 38,753 characters and repeatedly says TRYING_THE_MAXIMUM_URL.
Most users came from one regional online community, so this is not a census of the whole web. It is a record of what that community shared between 2009 and 2014.
I brought 0.mk back to test whether AI can now handle enough development, spam filtering, abuse review, monitoring, and support to make the service sustainable where the original economics failed.
Happy to answer questions about the crawl, the old data, or the rebuild.
hyperionultra 12 hours ago [-]
How did you managed to obtain that domain? Usually single digit or letter domains are “reserved”.
Intriguingly, the current Republic of North Macedonia (and its neighbors/constituent parts) has undergone a protracted and contentious naming dispute of its own.
Old web was kind of dumb anyway. You can put on the rose tinted glasses and feel elite about browsing some shitty site 20 years ago or enjoy the fruits of modern design.
prmoustache 8 hours ago [-]
> or enjoy the fruits of modern design.
Like web pages that load slower with high speed fiber than the old web did on 56k and ISDN connections?
inigyou 5 hours ago [-]
Hey now. He probably thinks that tokens do something, too. Be easy on him.
juleiie 2 hours ago [-]
If I want to read a book I have my e reader not a browser. Give me something good. hit my eyeballs with some pigments
This one is almost gone already.
Whoa, careful with your verb tenses! You just wrote a ghost story :)
its interesting because it sounds more correct to use have not yet because your life is still going on but its going on in the forward direction, not backwards in time to overlap an earlier HN state. So you could say its not right, but the sentence sounds/feels more right to preserve that your life is still going on.
We like that you are still here with us.
So when people construct versions of the old web today, they make a conscious decision to do something most people would not, and by doing so it sometimes comes off as indicating what they think of the other design decisions they chose not to make. It's not a bad thing, but it's not quite the same as it was anymore. That is why some of the spark is gone, in my eyes. It's kind of like counterculture being mistaken as the opposite of culture.
It is even more pronounced when people talk about "the old web" on old web websites all day.
TBH I do recall a lot of gloating over how AOL was full of clueless lusers who were inferior to the Internet
It being harder to join has the effect of their being a smaller more engaged user base. Less likely that there will be pressure from corporate interests, more likely that every user has a vested interest in both production and consumption.
edit: And of course, you'll see the same in dedicated niche forms and spaces like HAM radio. I imagine the folks TXing EME have a particular flavor that is well maintained.
Mind you, 200 years later, we still live in cities built for cars, not people. Rockefeller, Shell, BP and the like are still not dead, and still use everything they got to not go extinct while corrupting every politician that will accept their money. Just for a couple more years at a time, and a couple more, and a couple more...while destroying our planet like a cancer.
In 200 years we'll still have the web built for surveillance capitalism, not people.
There is no old web to go back to. Meanwhile state surveillance depends on its existence. Amazon, Apple, Google, Meta, Microsoft, three letter agencies, Palantir, and now the AI companies depend on it, and they're selling to nation states at a time. Again, it's just too juicy not to.
Back then we called them trusts, but now those big tech companies are seemingly worth more than actual countries. That's too much power without legislation, as they can effectively abandon the rule of law. Remember what happened with the Enron scandal? That was about billions, not trillions.
If we want a human web again, we need to build a new one, with privacy and secrecy first and not as an afterthought. No CA, and with mutual cryptography that acts like both a ledger and as a privacy shield for the users. No cookies, no tracking, no degenerate google controlling it all through owning the source code.
"Things used to be so good" is mostly rose-tinted glasses: What exactly WAS the "old web" anyway?
100s of half-baked sites hosted on Geocities, Yahoo, about pointless stuff? Epilepsy simulator websites splattered with gif-vomit? Everyone and their grandmother asking you to install their 10 toolbars?
Honest, serious question: What actually was of objective substance on the old internet that's nowhere to be found now?
You can find random pointless stuff now too, just that except Geocities/Yahoo it's Intsagram/TikTok/Twitter etc.
If you mean self-hosted websites, they're still here, and nothing's stopping you from creating your own.
If you loved all the Flash toons on Newgrounds etc there's unironically a lot more shorts and animations on YouTube or even Vimeo now, if you but search for them (the old gods like Weebl, David Firth are still there, and I suggest Sechi, DoodletmeGo, Nondescript Video Club, and からめる to start with then let the algorithm soak up the weirdness)
Just as the "internet" supplanted newspapers, magazines, and broadcast television for many people,
and how the newspapers replaced the town criers before them,
why shouldn't the "internet" be supplanted by a more accessible medium?
There's so much crap on the "old" and "middle age" internet: ads, scams, anal mods, trolls, bandwagons, vote wars, low-effort content and other repetitive noise that could be bypassed by just asking AI about something and make better use of your limited mortal lifetime, like watching that anime you've been putting off for years.
The old web was local forum created by gamers for gamers who hosted real life events. You could post asking for super specific advice to your locale. It was eaten by reddit.
There were cool IRC servers for all sorts of niche projects and topics. This still exists of course but a lot of it was eaten by Discord or Slack.
While I never participated, sites like LiveJournal provided both a unique outlet and view into other people's lives. I think a lot of social media like Instagram and Facebook eat this up.
I'm sure everyone has their own versions. And this is the old web that was gobbled up by mega sites which aim to commericalise the crap out of everything.
Btw I consider HN part of the old web.
But..what says something like Diaryland can't be brought back now?
The old web culture can't exist because advertisers won't let it exist on modern platforms.
Additionally the spread of information is significantly less organic because of relentless pushes to use algorithmic feeds along with far worse distractions from content than there used to be.
Again, what's preventing you from buying a domain name and putting up your own website? bypass the advertisers?
"Not enough people will see it", you say? So you want the reach provided by "modern platforms" but you don't want their rules?
On the old web, you would just stumble onto something deep but the barrier to get to it was fairly high. It could also just be total nonsense but at least you could see the passion behind it.
I agree with your general take. People talk about the old web as 'the good old days' but it really wasn't that great. I do have some fond memories of the mid 2000's when it was all starting to mature but also before corporate capture took hold. Even then, you would just end up stumbling around blindly most of the time.
There are so many piracy discords where you can't ask how to pirate something because that's against discord rules. I mean really? Why are you using it then?
and don't forget Big Billy's "no more tears" guarantee on all anal mods purchased from the official BBAM shop!
Pre-Google, specifically, pre-PageRank, and even more specifically, pre-SEO.
Your post is oddly hyper-focused on pointing out the negatives, probably due to the asshole hat you put on. There were lots of good sites with tons of good, free information, and because the main way people searched was using directories, we weren't inundated with 1,000s of results on the same topic, vying for a top-placement. The writing itself was much more laid back because SEO-optimization didn't exist. Also, those "half-baked GeoCities sites" were awesome! Of course there were a lot of junk ones, but again, it was just a bunch of people writing about themselves, their interests, and generally being creative. And again, NO ONE WAS TRYING TO MAKE MONEY, they were just doing it. No "like and subscribe." No gaming an algorithm. I used to love finding guitar tabs on those sites and I had a couple of sites with lots of tabs I did on my own for two different bands.
> There's so much crap on the "old" and "middle age" internet: ads, scams, anal mods, trolls, low-effort content, that could be bypassed by just asking AI about something and make better use of your limited mortal lifetime.
I feel like this was GP's point, no? Because people are just using AI for search now (this is the main reason I use it) that there will be no point in making these spamming sites anymore, which "opens up" the internet a bit for those of us who enjoyed the old way. Yes, it's all still out there, but it's kind of hard to find, and it would be kinda neat if more people discovered making their websites outside of walled gardens.
because people are hyper-focused on the negatives of the "current" web, or acting like the sites they used to love can't be recreated anymore..
> it was just a bunch of people writing about themselves, their interests, and generally being creative.
There still are. on YouTube, TikTok, Instagram, Reddit, 5ch and other Japanese subcommunities.. the fact that they're all "silo'ed" into these less-than-benign corpoctopuses is another matter, but people are still people'ing
Fair!
> There still are. on YouTube, TikTok, Instagram, Reddit, 5ch and other Japanese subcommunities.. the fact that they're all "silo'ed" into these less-than-benign corpoctopuses is another matter, but people are still people'ing
I know there are, but as per my comment, it's just harder to find and you used to almost not be able to avoid it. I don't think platforms like YouTube and Instagram are great examples due to the monotization... everyone is trying to "make it." But ya, I mean, I do know it's out there. But I think the whole point wasn't that this stuff doesn't still exist, just that maybe it will become the more prevalent thing again. At least that's how I read GP's comment. But who knows.
this comment brings up a good concrete example: https://news.ycombinator.com/item?id=49293553
Honestly I loved having domain specific forums that weren't beholden to a unified shareholder friendly content policy. Now they've been supplanted by a Reddit that's trying desperately to monetize content.
AI makes things less accessible because it centralizes power and increases the cost of publishing content. Anal mods is oddly specific, if you don't like butt stuff you could theoretically just ignore it. But okay. My time is better spent reading and understanding than waiting for some anti-bot crap to finish before I pick through largely irrelevant AI slop masquerading as an accurate summary.Those scams and that "crap" largely still exist on the "new" internet. Only now you're going to have to prove you're a human to view them. Where things have evolved is that instead of human written SEO drivel you've now got AI generated SEO drivel. It's to the point that I have to add "-deepwiki -ai" to pretty much every DDG search.
The problem is that AI companies are going full scorched earth. AI provides dog shit answers while ruining traditional search results. Or if you'd like a physical example they're physically destroying rare books to come up with poor, largely irrelevant AI drivel for the masses.
Today, if you own your own domain and fiddle with HTML then basically you're still part of the old web.
So basically, the old web required experts to help people get online, the new web negated them through homogenization.
The real experts from the old web basically became startup founders to monetize the homogenization and automation of their worldviews for the less tech-savvy (or specialized) folks.
People pine for the old web, but it was a churning, friggin' mess too. Search results were good though, but only when Google arrived. The standardization and ease-of-use from the new web has been a massive boon. I think things will change as people yearn for a wild-west, where there are new opportunities. The promise of a better life, seeing that the grass is greener, avoiding pain - those things always create golden ages.
Once AI favours people who can afford to pay for the best results.. or surveillance becomes too burdensome or dangerous, then watch this space.
The spirit of freedom always runs into trouble, when considering the notion of:
"Cut off the heads of others to make oneself appear taller."
Surveillance enables this. Decentralize again, but how?
ISPs - help or hindrance? Why/how can ISP charges be diverted to content creators?
But maybe that's more a measure of my own age and perceptions rather than an accurate representation of the various eras of the internet/web...
Two of my favorite sites [0][1] still online—Lurkers Guide to Babylon 5 and ex-astris-scientia—were started in ‘92 and ‘98 respectively.
Elegant content for a more civilized age.
[0] http://www.midwinter.com/lurk/
[1] https://www.ex-astris-scientia.org/
With that said, its domain has changed a few times iirc… so maybe not the best example given the nature of the article.
2009-2014 is Web 2.0, back when people still thought social media was a good idea.
PNGs? We had (mostly static) GIFs and we liked them.
Hit counters on the front page was mandatory of course.
I'd consider the facebook and mega era to be relatively new, the "old web" for me would be the one without centralization around a few giants, the era of random phpbb forums, private websites with "this site is under construction" banners and internet directories to find stuff.
https://britneyspears.ac/lasers.htm
Anyway Britney's 1998 debut album predates The Matrix, includes the song Email My Heart and came on "Enhanced CD" with tags and video. How about that
Now people don't give a shit about URLs. Webmasters casually move files around all the time because they feel like it, and if links get broken, who cares, that's the referring site's problem! Their beautiful file hierarchy is more important than the web staying connected!
That random 20 year old video of some middle schoolers doing a flying kick and breaking a vending machine? Yeah, just popped up on my feed yesterday.
0.mk, you had one job…
And now I wonder if Apache is even still common. I spent so much time fiddling with apache configs.
Also it's super easy to tell by looking at it, way too many LLM-isms. No need for a AI checker tool.
Did the zero ever mean anything? It's still a 3 character domain regardless if the first character is a zero.
I am not saying old content needs to be preserved forever, but so much content has factually been lost over time. Old logs from text-based MUDs for instance, even for MUDs that still exist today.
We hot-linked to all those image hosts because we couldn't imagine them disappearing.
Archive.org had incredible foresight and if it didn't already exist, I'd call such a project a pipe dream.
No, we hot-linked all those image hosts because we didn't want to pay to host it ourselves.
I enjoyed the emails I got from a couple people who were adamant I "hacked" their site because their "web developer" linked to images on my server... images that now said stuff like "I'm a loser bandwidth thief!", etc. (I never did use really nasty "shock" images-- just taunting stuff.)
This is a bit idealized. In practice copying data is not quite accurate (especially in bulk) and bit-rot is a very real phenomenon, both in flight and in storage.
You sometimes encounter it when dealing with files from the early '00s, it's very common to discover a few of them are corrupt, even if they've only ever been copied between harddrives.
It's copyright that causes anything to disappear from the web IMO -- torrents never die.
EDIT: I am aware that unseeded torrents do in fact die. But it really doesn't take much to seed a whole hard drive's worth of rarely requested data -- this also detects bitrot and so corrects errors automatically if you're not the only copy.
If you are, there's ECC, as well as making another copy.
There are many TinyMUD logs that were posted on Usenet, still to be found on Google Groups.
However, logging was controversial amongst mudders. It was almost always rude to log a private conversation without knowledge or consent; it was tacky to indiscriminately log while everyone was in the "hangout room" or Rec Room, as it were, and it was also bad form to post logs to Usenet or share them without redacting player names and other things.
But logging was built-in to most clients, and it was possible for server administrators to log (and hypothetically any malware-in-the-middle could log the cleartext, unencrypted TinyMUD TCP streams.) And many nefarious deeds by nasty players were exposed to the light when their logs were posted.
Let me introduce you to the alternatives: print media, film, stone engravings. That stuff tends to get burned and shattered and it takes FOREVER to make copies.
I'm being cheeky but I don't know what design change you could possible make to the web to make webpages not die.
I'd say it's one of the biggest design flaws of the current web, what with more and more content hidden behind paywalls, increasingly restricted WAFs, and rendered client-side via convoluted JavaScript.
Archiving and mirroring of old-style websites, delivered as static HTML, is simple and straightforward. 20 years from now, most web content from ~1996 to ~2015 will still be accessible, but much of today's web content probably won't.
Have an LLM “guess” random URLs seeded with words from a dictionary, iterating over each word and guessing a URL.
It guesses a lot of correct URLs. This is one method of “URL hunting” that doesn’t involve a 3rd party list or index.
Then just scan those pages for other URLs, visit them, and add a tally every time you come across a URL (for page rank).
Then search anything, see what the results are. You have invented a dark web search engine.
Guessing random words and .onion at the end probably won't get you anywhere, although onion is also quite a random word...
Also:
> Reply to any 0.mk email and the message lands in a feedback queue the AI reads, triages, and acts on
Is this dangerous? What about jailbreaking AIs and having it delete everyone's account?
The article doesn't seem to describe the cost of the AI solution. It does imply that it is lower than the cost of maintaining and supporting their service manually.
[1] https://wiki.archiveteam.org/index.php/URLTeam#cite_note-1 [2] https://lwn.net/Articles/683880/
0.mk started in 2009 as a passion project built by three of us. We worked on it for a few hours each week around our regular jobs. We eventually closed it in 2014 because the revenue (hint: no revenue) could not cover hosting, development, and the constant work of fighting spam and reviewing abuse.
The recovered historical corpus contains 657,607 links. For this analysis, we followed every one of them.
Of the 655,178 links with safe, crawlable targets, 76.7% no longer returned a loading page. After removing repeated destinations, 78.7% of the 492,620 distinct crawlable URLs still did not load. So duplicate links are not creating the result.
I use “did not load” rather than “gone” deliberately. Some URLs returned 403 or 429 and may have blocked the crawler. Pages that returned 2xx or 3xx count as loading even when they now lead to parked domains, login walls, or removed-content notices.
There is one large distortion in the yearly data. A single account created 83,398 URLs pointing to one hostname in 2011. At URL level, 92.5% of that year did not load. Count each hostname once and the result becomes 61.7%, almost identical to 2010 and 2012.
A few things I did not expect:
- 835 restored links point at Facebook’s old photo CDN. None loaded. - The first link ever shortened was a CSS stylesheet on a WordPress blog. - Someone shortened localhost on the second day. - The longest stored URL is 38,753 characters and repeatedly says TRYING_THE_MAXIMUM_URL.
Most users came from one regional online community, so this is not a census of the whole web. It is a record of what that community shared between 2009 and 2014.
I brought 0.mk back to test whether AI can now handle enough development, spam filtering, abuse review, monitoring, and support to make the service sustainable where the original economics failed.
Happy to answer questions about the crawl, the old data, or the rebuild.
“The name of the .mk domain consists of a minimum of 1 (one) and a maximum of 63 characters.”
Intriguingly, the current Republic of North Macedonia (and its neighbors/constituent parts) has undergone a protracted and contentious naming dispute of its own.
https://en.wikipedia.org/wiki/Macedonia_naming_dispute
Like web pages that load slower with high speed fiber than the old web did on 56k and ISDN connections?