{"id":648,"date":"2015-05-26T20:49:45","date_gmt":"2015-05-26T15:49:45","guid":{"rendered":"http:\/\/blogs.nd.edu\/emorgan\/?p=648"},"modified":"2015-06-01T22:59:01","modified_gmt":"2015-06-01T17:59:01","slug":"htrc-workset-browser","status":"publish","type":"post","link":"https:\/\/sites.nd.edu\/emorgan\/2015\/05\/htrc-workset-browser\/","title":{"rendered":"HathiTrust Resource Center Workset Browser"},"content":{"rendered":"<p>\nIn my copious spare time I have hacked together a thing I\u2019m calling the HathiTrust Research Center Workset Browser, a (fledgling) tool for doing \u201cdistant reading\u201d against corpora from the HathiTrust. [1]\n<\/p>\n<p>\nThe idea is to: 1) create, refine, or identify a HathiTrust Research Center workset of interest \u2014 your corpus, 2) feed the workset\u2019s rsync file to the Browser, 3) have the Browser download, index, and analyze the corpus, and 4) enable to reader to search, browse, and interact with the result of the analysis. With varying success, I have done this with a number of worksets ranging on topics from literature, philosophy, Rome, and cookery. The best working examples are the ones from Thoreau and Austen. [2, 3] The others are still buggy.\n<\/p>\n<p>\nAs a further example, the Browser can\/will create reports describing the corpus as a whole. This analysis includes the size of a corpus measured in pages as well as words, date ranges, word frequencies, and selected items of interest based on pre-set \u201cthemes\u201d \u2014 usage of color words, name of \u201cgreat\u201d authors, and a set of timeless ideas. [4] This report is based on more fundamental reports such as frequency tables, a \u201ccatalog\u201d, and lists of unique words. [5, 6, 7, 8]\n<\/p>\n<p>\n<a href=\"http:\/\/blogs.nd.edu\/emorgan\/files\/2015\/05\/catalog.png\"><img loading=\"lazy\" decoding=\"async\" src=\"http:\/\/blogs.nd.edu\/emorgan\/files\/2015\/05\/catalog-300x110.png\" alt=\"catalog\" width=\"300\" height=\"110\" class=\"alignright size-medium wp-image-652\" srcset=\"https:\/\/sites.nd.edu\/emorgan\/files\/2015\/05\/catalog-300x110.png 300w, https:\/\/sites.nd.edu\/emorgan\/files\/2015\/05\/catalog-1024x376.png 1024w, https:\/\/sites.nd.edu\/emorgan\/files\/2015\/05\/catalog-900x331.png 900w, https:\/\/sites.nd.edu\/emorgan\/files\/2015\/05\/catalog.png 1211w\" sizes=\"auto, (max-width: 300px) 100vw, 300px\" \/><\/a>\n<\/p>\n<p>\nThe whole thing is written in a combination of shell and Python scripts. It should run on just about any out-of-the-box Linux or Macintosh computer. Take a look at the code. [9] No special libraries needed. (\u201cFamous last words.\u201d) In its current state, it is very Unix-y. Everything is done from the command line. Lot\u2019s of plain text files and the exploitation of STDIN and STDOUT. Like a Renaissance cartoon, the Browser, in its current state, is only a sketch. Only later will a more full-bodied, Web-based interface be created.\n<\/p>\n<p>\nThe next steps are numerous and listed in no priority order: putting the whole thing on GitHub, outputting the reports in generic formats so other things can easily read them, improving the terminal-based search interface, implementing a Web-based search interface, writing advanced programs in R that chart and graph analysis, provide a means for comparing &amp; contrasting two or more items from a corpus, indexing the corpus with a (real) indexer such as Solr, writing a \u201ccookbook\u201d describing how to use the browser to to \u201ckewl\u201d things, making the metadata of corpora available as Linked Data, etc.\n<\/p>\n<p>\n&#8216;Want to give it a try? For a limited period of time, go to the HathiTrust Research Center Portal, create (refine or identify) a collection of personal interest, use the Algorithms tool to export the collection&#8217;s rsync file, and send the file to me. I will feed the rsync file to the Browser, and then send you the URL pointing to the results. [10] Let\u2019s see what happens.\n<\/p>\n<p>\nFun with public domain content, text mining, and the definition of librarianship.\n<\/p>\n<h2>Links<\/h2>\n<ol>\n<li>HTRC Workset Browser &#8211; <a href=\"http:\/\/bit.ly\/workset-browser\">http:\/\/bit.ly\/workset-browser<\/a><\/li>\n<li>Thoreau &#8211; <a href=\"http:\/\/bit.ly\/browser-thoreau\">http:\/\/bit.ly\/browser-thoreau<\/a><\/li>\n<li>Austen &#8211; <a href=\"http:\/\/bit.ly\/browser-austen\">http:\/\/bit.ly\/browser-austen<\/a><\/li>\n<li>Thoreau report &#8211; <a href=\"http:\/\/bit.ly\/browser-thoreau-about\">http:\/\/bit.ly\/browser-thoreau-about<\/a><\/li>\n<li>Thoreau dictionary (frequency list) &#8211; <a href=\"http:\/\/bit.ly\/thoreau-dictionary\">http:\/\/bit.ly\/thoreau-dictionary<\/a><\/li>\n<li>usage of color words in Thoreau \u2014 <a href=\"http:\/\/bit.ly\/thoreau-colors\">http:\/\/bit.ly\/thoreau-colors<\/a><\/li>\n<li>unique words in the corpus &#8211; <a href=\"http:\/\/bit.ly\/thoreau-unique\">http:\/\/bit.ly\/thoreau-unique<\/a><\/li>\n<li>Thoreau \u201ccatalog\u201d \u2014 <a href=\"http:\/\/bit.ly\/thoreau-catalog\">http:\/\/bit.ly\/thoreau-catalog<\/a><\/li>\n<li>source code &#8211; <a href=\"http:\/\/ntrda.me\/1Q8pPoI\">http:\/\/ntrda.me\/1Q8pPoI<\/a><\/li>\n<li>HathiTrust Research Center Portal &#8211; <a href=\"https:\/\/sharc.hathitrust.org\">https:\/\/sharc.hathitrust.org<\/a><\/li>\n<\/ol>\n","protected":false},"excerpt":{"rendered":"<p>In my copious spare time I have hacked together a thing I\u2019m calling the HathiTrust Research Center Workset Browser, a (fledgling) tool for doing \u201cdistant reading\u201d against corpora from the HathiTrust. [1] The idea is to: 1) create, refine, or identify a HathiTrust Research Center workset of interest \u2014 your corpus, 2) feed the workset\u2019s [&hellip;]<\/p>\n","protected":false},"author":92,"featured_media":0,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-648","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"_links":{"self":[{"href":"https:\/\/sites.nd.edu\/emorgan\/wp-json\/wp\/v2\/posts\/648","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/sites.nd.edu\/emorgan\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/sites.nd.edu\/emorgan\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/sites.nd.edu\/emorgan\/wp-json\/wp\/v2\/users\/92"}],"replies":[{"embeddable":true,"href":"https:\/\/sites.nd.edu\/emorgan\/wp-json\/wp\/v2\/comments?post=648"}],"version-history":[{"count":5,"href":"https:\/\/sites.nd.edu\/emorgan\/wp-json\/wp\/v2\/posts\/648\/revisions"}],"predecessor-version":[{"id":655,"href":"https:\/\/sites.nd.edu\/emorgan\/wp-json\/wp\/v2\/posts\/648\/revisions\/655"}],"wp:attachment":[{"href":"https:\/\/sites.nd.edu\/emorgan\/wp-json\/wp\/v2\/media?parent=648"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/sites.nd.edu\/emorgan\/wp-json\/wp\/v2\/categories?post=648"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/sites.nd.edu\/emorgan\/wp-json\/wp\/v2\/tags?post=648"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}