I think people, possibly including me, get irked with Demand Media et al more because they're more successful than we think they deserve to be rather than because they actually decrease the value of the SERPs. For SERPS where DM ranks well, the results prior to DM existing generally pretty much sucked. Maybe that is a Google issue, maybe that is an Internet issue (memo to Internet: middle aged women exist, please write for them, kthxbye), but for whatever reason, if you routinely Googled for [how do i make a blueberry pie] every week for the last ten years I don't think you ever had an awesome search experience.
DM pages are adequate for much of what they rank for, in much the way that USA Today is an adequate newspaper, your local state school provides adequate degrees in history, etc etc. They're adequate in a scalable manner, though, and they understand Google much better than the average publisher, which means they get visibility in excess of what some people might expect.
If I wanted to bake a blueberry pie, I'd go for that second page every day of the week, but it is highly non-obvious to me that it is a better result qua search engine result than the DM page. I love this example because I think Google fundamentally doesn't think [how do i make a blueberry pie] is looking for a blueberry pie recipe. Most searches will not actually convert to pies. For the 98% of searchers who merely want to satisfy their pie voyeurism need, the DM content may well be better.
(memo to Internet: middle aged women exist, please write for them, kthxbye)
Many years ago, when I was working at Joseph Beth Bookstore (where I suspect Borders stole a lot of their ops manual ideas from) management there told me that the primary demographic there was middle aged women. Middle aged women often have disposable income. Many of them are married homemakers and have tons of spare time. I suspect much of Oprah's success is built on that demographic.
If I were Oprah, I'd position her cable channel as just a part of her latest "new media" venture -- just one outlet for Oprah branded content. In her shoes, I'd make diverse investments in mobile and in TV-integrated platforms. I'd offer advertisers packages for reaching all those different outlets simultaneously. If Oprah does this, I'd recommend working with her. If she can't see this path clearly, then I'd recommend disrupting her.
My mother-in-law used to be an Oprah fan but she recently 'cut the cord' with the cable company. Personally I think people like her could really enjoy an experience like Digg or Reddit if all the stars aligned correctly.
Maybe something like reddit, but presented through things like Google TV and Apple TV. If someone can morph social media viewing/browsing into something that emulates channel surfing, then that idea will easily be worth 100's of millions.
EDIT IDEA: If Oprah doesn't have someone doing the following for her, then some group with the right CV items should hop on it! Start finding or putting key moments from shows followed by middle aged-women up on YouTube. (Oprah will make up a lot of this content.) At the same time, launch a social media site tailored for middle-aged women and create APIs so that it can be integrated with TV convergence devices.
Not only will Oprah make more money on her own content this way, she will become the middlewoman for a big chunk of the other content aimed at this demographic.
I don't know what a 'SERP' is but I don't think anyone shows up at Google bright and early in the morning with the burning desire to deliver the 'USA Today' quality of search. In fact, we had it before Google came along and it left much to be desired.
I don't know much about blueberry pie searches either or if the quality of Google results is really in decline. It seems pretty reasonable to expect original content (StackOverflow) to show up before copies of the same content. Or the top search engine to aim for a quality standard above 'USA Today'. Otherwise we can just use USA Today instead of Google.
The StackOverflow scrapers ranking higher than SO is the thing that most irks me about the big G at the moment. That and all the sites like wareseeker that just act as pointless aggregators of FLOSS download sites or forum sites.
Last time I tried Bing though they weren't any better, and DuckDuckGo had really low coverage. Maybe it's time to try again.
I use DuckDuckGo as my main search engine, since I figured out the bang syntax. If I don't like the results, I just resend the same query with !g prefixed (!google works, too, and so does !bing).
I think I see where you're going, but I disagree. To me, link #2 is superior even if I just want pie voyeurism. (mainly, it has instructional pictures) I would click on #2 every time.
However, if #2 is buried on the 3rd page of results like it sometimes is, then I would not....
There is however, a definite gradation of google spam. ehow is somewhere near the top... domain squatters somewhere near the bottom...
Picking Virtuous over Demand Media is your opinion.
Personally, I've been using eHow's results for cooking for about half a year now.
I don't need pictorial aids for cooking--I know how to cook. What I need is a recipe, and a series of concise steps to follow. It's easy to toss it up on the laptop or table, then glance at every so often to make sure you haven't strayed too far from the recipee.
Demand Media is like a virtual cookbook, Virtuous is like calling your parents for help and having them hold your hand through every step.
You know very well that this is not how search works. You don't "ask a question" to the search engine like you would ask your grandmother.
You type in words that you expect to be in the pages you're looking for, and the search engine lists pages that actually contain ALL of those words.
One of the main improvements of Google in the very early days was that it used the AND operand by default, whereas competing search engines used OR by default, resulting in an incredible amount of noise.
In essence, searching for "how do i make a blueberry pie" (with quotes) should return only spam, because only spammy and SEO optimized sites would contain the phrase as such. A real recipe would maybe contain the phrase "how TO make a blueberry pie" but not "how do i..."
- - -
I think your point was that "middle aged women" don't know any of this.
It would be arguable (probably wrong, but still) that people who don't know this, who didn't make the effort to understand a little how all of this works, deserve the spam they get.
There is a good way to discriminate between good and bad content, and that is to know a little about what you're searching in order to search for words that will be present in good quality content and NOT in spammy pages.
For example, it's reasonnable to expect a good recipe to give instructions in metric system as well as imperial; if you add "celsius" to the search then the second (informative) recipe arrives first:
I think your point was that "middle aged women" don't know any of this.
No. That is overbroad, untrue, and would be very injurious to my professional reputation. I said that the Internet is skewed away from producing content responsive to their needs, which is about as controversial as saying that they are slightly underrepresented on HN relative to, I don't know, twenty-something males.
Non-technical users frequently use natural language search. The experience for natural language search is fairly poor. There are many classes of search which offer poor experience, but it is the one which leaps to my head first because I deal with non-technical users every day.
people who don't know this, who didn't make the effort to understand a little how all of this works, deserve the spam they get.
Words cannot express the depth of my distaste for this position. I will accept that "Google screwed up" or "I screwed up" if one of my users has a suboptimal Internet experience (which starts at Google because Google is the Internet and ideally ends at my site), but I cannot accept that she is responsible if she has a poor user experience. We've got the teams of PhDs, the highly paid SEO consultants, and the lifetime of building an accurate mental model of how the devil box works. She wants to teach kids to read, not learn magic incantations. It should -- "should" in the sense of "would be optimal for the business", "would be optimal for society", and "as a moral imperative for computing professionals" -- just work for her.
(I really don't understand your first/second sentence (what's injurious?) Also, I'm 40.)
> I will accept that "Google screwed up" (...) if one of my users has a suboptimal Internet experience
I respect that, and from a business point of view you're very right.
The problem is, how do you improve her experience without screwing up mine? Why can't I search for pages that actually contain all the words I'm looking for, as I typed them, and not words "that were present in the page linking to this page" or words Google think I want although I didn't type them in?
From a "moral" point of view (which you brought up), if she "wants to teach kids to read" maybe she could start by learning how to spell?
Most of the world does not think like a nerd. Not even remotely. Remember that when working.
Clue: remember all those things that many of us nerds thought the iPad really needed?
Corollary: why is WinARM likely in deep trouble? Because it's three years late to market, and Microsoft is only recently shown themselves capable of designing a UI for mortals, and the question is whether the Windows UI or the (ill-branded) Windows Phone UI will win the internecine politics. (And I'm betting on the mass of the Windows group.)
Implementation detail: your Google Internet searches are already tuned, if you're logged into Google.
And for completeness, when using the phrase "the problem is..." remember too that with problems arise opportunities, and through opportunities can arise profits.
I don't know why the above comment is being downvoted, but I'm guessing it's because it sounds "elitist" (which is apparently a very great crime).
To elaborate, then: I agree with the parent comment that it's Google's job to make everyone's experience optimal (and not the user's), and it's certainly in the best interests of Google (or any business) to cater to the needs of as many of its customers as possible (although in the case of Google, as has been pointed out many times before, users are in fact the product).
But I would argue that the real elitists are people who think "middle aged women" shouldn't be expected to actually learn how to use machines.
"Middle aged women" (why single them out?) use machines all the time, whether at work or at home. They're expected to know how to use a spreadsheet, a word processor, a food processor. And they do. But somehow this expectation is lifted for "the Internet". Why?
A search engine is not a person; it's certainly not a mind reader. A search engine is just a machine.
Plus a few other pre-eminent and trusted food-networks (Allrecipes) - searching for "Blueberry Pie" on each one, I'd then scan the quality of the comments - looking for insight into other chefs (cross checking their history to see if they, in turn, can be trusted) who clearly have tried out the recipes, and have made relevant comments. I would then identify the recipe that looked most likely to work for me.
I would expect no less from a sufficiently advanced search engine in this, and all other domains.
The Food Network? Really? The home of "semi-homemade"...? ;-)
About Clarke's law, here's an observation by George Bernard Shaw: "Build a system that even a fool can use, and only a fool will want to use it."
Quotations aside, the process you're describing is certainly excellent; it's probably what Blekko is trying to pull, in a scalable way. It'll be interesting to watch how it plays out.
The second iteration, of course, is to engage with every (valid, trusted, revenue generating, etc...) customer who searched, determine the quality of the results, and then feed _that_ information back into the algorithms. You could then bias based on domain experts (world class chef's feedback on BlueBerry Pies more important than an anonymous user)
It may be the case that (AllRecipes, TheFoodNetwork, Etc...) are NOT the best place to search for a recipe, and that, indeed, "http://pickyourown.org is knocking it out of the park this week.
There is a lot of room for search to improve - I think the company that beats google (if it's not google that does so first), will be the one that manages to start creating the search<-->Consumer<-->Search feedback quality looop.
Google may already collect enough data to do this. They track clicks on the search results, so they can see whether you liked the results, whether you went back to a different result after visiting your first, and whether you modify your search terms for another search because the first one did not work out.
1) If you don't understand tax law, should you have to pay more tax? If tax law is simplified (so everybody pays a "fair" amount), why should people who have bothered to structure their affairs in a tax-efficient manner lose out?
2) Google has an "advanced" tab. If they really wanted to, they could have a "fuzzy match" section, and a "required literal" section. They could also do some funky stuff with leximes, but it's just not worth it, even for advanced users.
Non-technical users frequently use natural language search. The experience for natural language search is fairly poor.
Spot-on.
When my wife has trouble finding something and asks for help, this is usually the issue. The best approach isn't to ask Google a question. It's to picture in your mind what your target page looks like, and enter that into the search query.
But this just underscores what all of us here already know: natural language processing is hard.
"You don't "ask a question" to the search engine like you would ask your grandmother."
It doesn't matter one jot 'how search works'. What matters is how the majority of their users think it works. And they think it works by asking a question. So publishers need to work with that.
You're right, and yet I disagree with you (which I guess means I'm wrong).
Your position is the path of least resistance: indulge users in their ignorance. It's optimized for the short term, and in other contexts leads to great catastrophes (fast food, for example).
I wish I could find a link to an article I was reading a while back, which pointed out that most users who actually use search engines beyond typing 'facebook login' every morning work out how to search within a couple of weeks at most. Given that, it seems like a bad idea to optimise for the first week at the expense of every day for the next few years.
Sadly, I have no idea where I found that and don't recall if they had any factual data to back it up :(.
I think your point was that "middle aged women" don't know any of this.
No. The web (and computers) should adapt to us, humans. Not the other way around. If people want to search for a question, then that's the correct way to search.
The reason 'spammy' sites show up high for that query is because they are better at knowing what people are searching for.
But that is not true of any other human activity. Every human activity is learned.
The correct way to eat is not to throw handfuls of food towards your face, hoping that some will end in your mouth. That is how toddlers eat, until they are taught otherwise.
The correct way to write is not to scribble on a table with the wrong end of the pen. Etc.
Yet you argue that the correct way to use a search engine should be "how you do it when you first try it and don't know anything about it". This is inconsistent with what you have been doing since you were born.
But that is not true of any other human activity. Every human activity is learned.
And most of those learned activities are eventually replaced by better ones. People learned that the right way to portion food was by tearing with their hands until someone invented the knife. People learned that the right way to draw was with a stick in the dirt until someone invented paper.
What you think is the right way to do anything is only that because you haven't discovered or been taught a better way to do it, and that includes searching the internet.
Yet you argue that the correct way to use a search engine should be "how you do it when you first try it and don't know anything about it".
Actually, I think the argument is that a good way to be successful in business is by offering your customers more perceived value than your competitors. In search engines, it's more valuable for your users to not have to learn how your algorithm works in order to craft a search query that will satisfy their needs. It's valuable to be able to use natural language to search.
Right or wrong, the business that refuses to provide that value because they disagree with it is going to fail miserably.
I totally, completely and unconditionaly agree with your last point: it is absolutely in the best interests of Google to respond to every query with relevant results, however it's formulated.
If 99% of Google users ask Google questions like they would a person, then Google should be able to provide answers; that is true even if only 10% of users used it that way.
My point however is not about Google: it is about the user. The user would gain from learning how search engines work, and formulate their queries accordingly.
It's often stated that users shouldn't be bothered to learn how to use your service, because "they have more important things to do" and "there are other services to choose from".
This is good advice for businesses (common sense, really); it is bad business to expect users to dedicate time and effort to use your system.
But I have a hard time accepting that this is an excuse for users to ever learn anything. At the same time as it is Google's responsibility to serve users as best it can, it is each user's responsibility to try to improve their mastery of such tools.
> In search engines, it's more valuable for your users to not have to learn how your algorithm works in order to craft a search query that will satisfy their needs.
Well, yes. But in this case it's not an algorithm. You just ask it for documents containing your search terms. It's about as straightforward as it gets. It's not computer programming, it's not even doing long divisions or whatever.
You can't seriously assume that method of searching is going to completely stump anyone.
> It's valuable to be able to use natural language to search.
Maybe you're forgetting here that not the entire world speaks English. Unless your natural language search engine is able to understand nearly every language in the world, you're still forcing the users to formulate a query in English.
And I'm willing to argue that for, say, your average middle-aged German, formulating a query in English is going to pose a much bigger problem than doing a conjunctive term search, as the latter carries transparently to most languages in the world. (not all of them--I should check which ones btw--but even in those cases it's not nearly as difficult to find a fix for that than it would be to somehow port your brilliant semantic context analysis engine to be able to speak yet another completely different language).
Give it a couple of decades, maybe we'll have usable NLP then, but before that time, I am convinced that even for the layman, searching for documents to contain certain terms is a lot more straightforward, likely to yield desired results and in quite a few ways actually easier to use than what has to pass for NLP currently.
The examples you cite are mostly about physical actions (e.g. eating), your advice to adapt yourself to the world is sound because we can't change how the world and physics and matter work. However a search engine is a totally non-physical thing. Software has no body and essentially no limits like a pen does. We can make software do anything we want (almost). We can try to make software that understand how people ask questions. That's what we should do.
A better example would be languages. They are entirely intellectual and change all the time. Don't like speaking a certain way? Then change it, and it might take off! No reason to limit us all to Latin, we've invented all the other languages.
> We can make software do anything we want (almost).
We can build new software to do anything we want (if we know how to program). But some individual somewhere cannot have Google behave in some specific way. In that sense Google is very much a physical object of the world.
> We can build new software to do anything we want ...
I'm being really pedantic here, but you actually can't. There are actually uncountably which are impossible to solve with an algorithm. There's a (by necessity) incomplete list of examples at: [1].
I'd agree with you, if we were talking about something that is even slightly complicated. But it's not. Searching with the old-school Google syntax is asking the question "what documents contain these words X Y Z?".
It's not that hard, in fact very easy, to wrap your head around that. It's easier than looking up a phonenumber in the phonebook, or a business in the Yellow Pages. It's even easier than using the term index in the back of a book.
While asking a question might be easier than that still, it also has disadvantages. It's impossible to ask an exact question on a very specific subject. Try it in real life. You'll find you need context, or at least a series of back and forth questions to arrive at the answer you want to have.
If a search engine would implement the same method, any advantage from using questions over a conjunction of terms is negated. People would get tired of typing long question phrases, and unless the natural language processing of the search engine is absolutely perfect, they'd get pretty frustrated quickly because the search engine would interpret some of the questions wrongly.
For simple queries like "where can i find a recipe for blueberry pie?" or "what is the capital of Denmark?", we already have seen search engines can do this pretty well. For anything more complicated, they fail sometimes, if not most of the time.
It's not that hard for people to search "recipe blueberry pie" or "capital Denmark" as you think, either. People can get used to barking commands like this quite easily. Look at Star Trek ("earl grey, hot") or many other scifi series featuring a semi-intelligent computer. People almost expect it to communicate in terse phrases like that.
The big advantage, also for the layman, is that a conjunctive term query will nearly always yield the result they were expecting, because it doesn't leave much room for ambiguity. And if it does, it's when query terms hit a range of documents that weren't intended, but still match, again I refer to scifi, as well as fantasy, it may be frustrating, but it's also somewhat endearing (putting the user in a position of superiority) in the sense of an intelligent robot taking a request too literally. Or a genie in a bottle. Or an ET alien. Or commander Data. Or whatever.
On the other hand, if the computer interprets a full question in the wrong way, people get more annoyed, because the "too literal" explanation doesn't work here, after all, you typed a complete sentence, it should be obvious what you meant, right? So instead, people get the idea that the computer is simply not listening. And that makes it a lot harder to come up with a "better" question that would give them the answer they need.
Now, of course, as soon as search engines are able to do perfect language processing, as well as guessing a whole lot of additional clues from whatever bits of context they can get, typing questions into a search engine might be the better and easiest way. But natural language processing isn't really going there any time soon. Just look at the search engines that try this, how well they are doing now compared to how well they were doing 5 years ago, and there's not much progress made.
On the other hand, in the area of text-query processing, organising of data, cataloguing data, we have made tons and tons of progress. There's tagging, social network friends recommendation, and vastly superior methods of indexing large amounts of text documents for logical queries.
> You type in words that you expect to be in the pages you're looking for, and the search engine lists pages that actually contain ALL of those words.
It's not as simple as that though is it? Otherwise how would those GoogleBombs or whatever they're called work where searching for 'warmongering idiot' or something turn up George Bush's biography on whitehouse.gov. I think search result quality is a bit more involved than just testing whether a page contains a set of words.
> I think search result quality is a bit more involved than just testing whether a page contains a set of words
Google bombs work because Google includes in the words of the current page, words in the links leading to the page (inside the anchor), and even, I believe, words in the page linking to the current page, outside of the link itself.
And then fields are weighted to calculate relevance. Those long, long urls with the title of the page in the url started to appear and proliferate when people noticed that words in the url were given an important boost factor by Google.
This broke the Internet a little, by the way. That's why we need links shorteners now, with urls containing whole goddamn paragraphs.
I am reasonably search-savvy, and when I am starting to research something out of my comfort zone, I often start with a "how do I…" search. I sometimes find exactly what I need with it, but more often I get enough information to perform other, better searches.
Your assumption that people don't search that way is dead wrong, though: I've seen many non-technical people do just that sort of search. My mother being one.
I know: the plural of anecdote is not data and all, but I'm fairly certain most people don't think of search the way you suggest.
On the other hand it took google more than a couple of weeks to get rid of kods.net, too. Where ehow is readable, kods was, as far as I could tell, just a lot of computer generated oracle words mixed from other sources.
DM pages are adequate for much of what they rank for, in much the way that USA Today is an adequate newspaper, your local state school provides adequate degrees in history, etc etc. They're adequate in a scalable manner, though, and they understand Google much better than the average publisher, which means they get visibility in excess of what some people might expect.
P.S.
Demand Media: http://www.ehow.com/how_2933_make-blueberry-pie.html
Virtuous publishers on the Internet: http://www.pickyourown.org/blueberrypie.php
If I wanted to bake a blueberry pie, I'd go for that second page every day of the week, but it is highly non-obvious to me that it is a better result qua search engine result than the DM page. I love this example because I think Google fundamentally doesn't think [how do i make a blueberry pie] is looking for a blueberry pie recipe. Most searches will not actually convert to pies. For the 98% of searchers who merely want to satisfy their pie voyeurism need, the DM content may well be better.