Clusters keywords by result-page overlap instead of by string similarity. Two phrasings that return the same pages are one page's job; two that read alike but return different pages are two. Text-similarity clustering gets this backwards and produces pages that compete with each