28.04.2022
Vortragsreihe: ZIH-KolloquiumZIH-Kolloquium: "Web Archive Analytics"
The talk recaps how the Web has been systematically archived since around the turn of the millenium, and how popular resources are being exploited at ever increasing scales until today. Besides commercial players such as Google or Microsoft, whose search engines and many other products depend on Web archives maintained in-house, many non-profit organizations are archiving the Web as well. Chief among them is the Internet Archive, which was founded around 1996, alongside Google, and which has meanwhile grown into a fully-fledged digital library with the goal of providing "universal access to all knowledge". In collaboration with the Internet Archive, the partners Bauhaus-Universität Weimar, Leipzig University and Martin-Luther-Universtät Halle have downloaded a substantial part of the the Internet Archive's Web archive, which forms the basis for many joint research projects with applications in information retrieval and natural language processing.
Martin Potthast is head of the Text Mining and Retrieval Group at Leipzig University. His research focuses on information processing, the challenges arising from assessing the credibility and originality of information obtained online, and the development of information systems. Martin Potthast has contributed to the fields of information retrieval, natural language processing, and data science. Several of his achievements have been awarded with scientific prizes.
The colloquium is free of charge. Language: English
For joining this lecture please use the following link:
- For participants with ZIH login: Link ZIH-Colloquium
- For participants without university login: Link ZIH-Colloquium