Skip to content

Comment on Robots.txt for the NYT has a specific exclusion for an 1996 news article

Comments

I've added a special exclusion to a robot.txt file for a specific article. It was some years ago while in college. The article in question was about the presentation of an assisting professor who had some kind of misunderstanding with the campus newspaper and therefore the article wasn't especially positive in tone. Couple of years later I was the sys admin of the newspaper website and a letter arrived in my university email. The professor had found that I'm responsible for the website and had sent me a tearful story about how this article is ruining her life, because it is the top Google result for her name, and how she had spent thousands of dollars on scammers who had promised to change that, and she was asking me to remove the piece. Long story short, I forwarded the case to the newspaper editor at the time and she agreed to let me add a line to the robot.txt.

Edit: newspaper -> newspaper editor

Out of curiosity, I checked the student's newspaper website. It turns out that they've made redesign after my time and they've removed the robots.txt file. However, googling the name of the professor in question returns much more recent results and the article is hard to find. It turns out that Google's algorithm buries some stories over time. One more decade and noone will be able to find this part of her history unless they know what they are looking for.

I find that Google discriminates against old content just because.

Old, to the point, websites hand-written in notepad.exe are rarely in the top 5-10, even when they have precisely the answer you’re looking for.

It's incredibly hard to find any information that wasn't created this year.

Not sure if it's Google discriminating against old sites, or if the new sites just have such incredible levels of SEO-fu that a static, just-contains-what-you-want site has no hope of getting selected for results

Except when it isn't. I don't know if I'm uniquely bad at web search these days, but Google seems to do the exact opposite of what I want. I.e. when I need some solid information about a piece of technology or knowledge, it'll spam me with content marketing bullshit published last week. But when I need something recent - like a current opinion on some software alternatives, or installation instructions, all I get are results from 5+ years ago.

I don’t know if we’re both in that case and Google has simply moved away from how we learned to coerce it (and other search engines) over the years, our habits now yielding negative returns unbeknownst to us, but I find myself in a similar situation.

Month after month, year after year, search results get worse, and harder to sift through. At this point, google is mostly a way to find wiki articles despite typos (because Wikipedia’s search is bad and can’t handle typos).

Google has definitely changed in the past few years. I've noticed that it ignores quotes when it feels like it, and is very aggressive at "synonymous" substitutions (which in many cases aren't - e.g. I've seen it replace "FreeBSD" with "Linux"). It's making it harder and harder to make queries for something specific.

I thought I am the only one! Last week, I was trying to find something with the quotes and it completely ignore it. It took me a few queries to find what I am looking for.

Use verbatim, as per my comment up thread.

<outlandish conspiracy theory>

Google does this to make every tech worker not working for Google, and thus Google's competition, less productive.

</outlandish conspiracy theory>

Or maybe the google AI has a funny side. Right now it's probably laughing its ass off reading teMPOral's comment.

I use duckduckgo and it always finds me PostgreSQL documentation for version 9.3 or 8.x or something obsolete like that during general search.

Google gave me Python 2 results for every doc query for a looooooonnnnnngggg time, I just clicked to the newest version when I got there.

That was mostly PSF's fault since they had the default documentation URLs setup to point at the Python 2 documentation for way longer than they should have...

Oh really? Interesting. I always thought it was because Google never switched to Python 3.

Ha, me too, but then I've also been caught out on 'latest' (which I typically reflexively click on after opening) having some crucial difference from what we're using (12 vs. 13) too, so hard to win there.

I don't know off-hand if it's possible with the search (e.g. it might use a last-selected version cookie) but we should probably be using version-specific 'bangs'. `!pg12` or whatever, in my case.

So does Google, and it’s even worse there. In Google, the first page for "postgres window functions" is version 9.1, version 9.3, and 8 random tutorials. DDG has 9.1, 9.3, tutorial (the same as Google…), but then you at least get 12, 11, and then some more tutorials.

Or it could also be that people are more often clicking on recent links or searching for "something something 2021". I know I usually give higher preference to links with more recent date attached as that just-contains-what-you-want site from 2010 will potentially be outdated by now. Of course it depends on subject of your search, but id imagine it's true more often than not in big scale.

This is the thing - for some things (e.g. searching for tech to use), you care about recency. But for something you absolutely don't. But the search engines magic seems to throw this all into one bucket and notice that people are more clicking the '2021' links. And the consequence of this is SEO optimization done by automatically updating all links to current year - I searched for something on 1st Jan and already got a bunch of '2021 comparisons'

It's a common content marketing tactic to update post titles to include "for 2021" a few weeks ahead of the new year. It's a tweak to get it ranking, as every other content marketer is doing the same.

Google often delists articles over around 15 to 20 years after they are put on the internet.

Google works for the masses. It assumes synonyms and aliases are great. For example, searching for phrases like debian <something>, often bring ubuntu help pages to the top results.

Worse if you are looking for Dave Smith. You'll get David Smith, because Dave=David, even though you know Dave never identifies as that.

Then add it fixing spelling errors. Which often are not. Rare term? Must be a spelling error!

Part of this, is because most people are non-precise, and also because Google wants voice input to work well. So there, their, they're are the same thing, but can also synonym to things like them, and they.

Point is, Google is trashy for any search not involving cats, or explosions. Quotes barely work, so the only way to mitigate this a bit, as google removed the + search modifier a decade ago, is always, use verbatim, under search tools.

And hope Google hasn't broken it that day. Which they do often.

I find that Google discriminates against old content just because.

Information rots over time, decaying from true to false. Some facts are eternal, many aren't.

Similar thing here. I wrote a blog post about an online "contest" that was actually a bit of a scam. For years afterward I'd periodically get email from the person asking to take the blog post down because its prominence in search results was souring job prospects. I eventually relented, not because I cared about him but because he mentioned that he had started a family and it was affecting them as well. I didn't want to hurt innocent people. I didn't want to censor my blog either, but I did add a robots.txt so it wouldn't show up in search results.

Since then it seems he really has gotten it together, and even had a project show up on the first page here. So now I suppose the robots.txt entry doesn't matter much, but it's still there anyway.

This is now institutionalised as The Right To be Forgotten. Every search engine doing business in the EU has implemented it. Most probably only offer it to Europeans.

The professor is not an EU citizen and the right to be forgotten does not cover unflattering articles in student newspapers.

It can cover exactly that scenario depending on the situation.

The right is for removal of search results from particular queries on Google etc.

The newspaper doesn't have to do anything.

Precisely.

I believe it not only covers, but is intended to deal exactly with that - news pieces about individuals that are no longer relevant, but the individuals want to get out of search indices.

The rights to privacy and to be forgotten could be superseded in the cases of significant public interest. Especially the right to be forgotten is in regard to petty crimes, revenge porn or stupid posts done by minors. Right to be forgotten does not allow for censorship or rewriting history at the will of the subject of an article.

Something that Wikipedia reminded me, rtf is obligatory only in the EU and search engines are not obliged to comply with it on its international websites.

https://en.wikipedia.org/wiki/Right_to_be_forgotten

That is one lucky professor -- how did she get granted an exception? There are tons of unlucky people who suffer as a result of bad/shoddy/vague reporting, how do they also get a carve-out?

In my case, nobody was actually malicious. It was a stupid situation caused by the stupid actions of both sides from what I've been able to discover. Nobody meant to cause harm but to tell their own story. I tried to find a compromise that would work for everyone and it worked this time.

Like many people who get exceptions. They asked.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.