CERIAS Weekly Security Seminar – Purdue University
Commercial web sites are more dependant than ever on being placed prominently within the result pages returned by a search engine to be successful. "Spam" web pages are web pages that are created for the sole purpose of misleading search engines and misdirecting traffic to target sites. Certain classes of spam pages, in particular those that are machine-generated, diverge in some of their properties from the properties of web pages in general. As a result, these pages can be identified through statistical analysis. We have examined a variety of such properties, including linkage structure, page content, and page evolution, and have found that outliers in the statistical distributions of these properties are predominantly caused by web spam. Joint work with Mark Manasse and Marc Najork. About the speaker: Dennis Fetterly is a Technologist in Microsoft Research\'s Silicon Valley lab, which he joined in May, 2003. His research interests include a wide variety of web related topics including web crawling, the evolution and clustering of pages on the web, and identifying spam web pages.
En liten tjänst av I'm With Friends. Finns även på engelska.