{"id":1052,"date":"2022-04-01T18:15:23","date_gmt":"2022-04-01T13:15:23","guid":{"rendered":"https:\/\/sites.nd.edu\/emorgan\/?p=1052"},"modified":"2022-04-01T18:24:21","modified_gmt":"2022-04-01T13:24:21","slug":"istl","status":"publish","type":"post","link":"https:\/\/sites.nd.edu\/emorgan\/2022\/04\/istl\/","title":{"rendered":"Issues in Science and Technology Librarianship"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">I asked a colleague (Parker Ladwig) if there was a blog he thought might be worthy of archiving, and he mentioned <cite><a href=\"https:\/\/journals.library.ualberta.ca\/istl\/index.php\/istl\">Issues in Science and Technology Librarianship<\/a><\/cite> (ISTL). I took it upon myself to see what I could do.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">After looking more closely at the site, I guessed the underlying technology was not blog technology but rooted in the venerable <a href=\"https:\/\/pkp.sfu.ca\/ojs\/\">OJS journal publishing system<\/a>, and OJS robustly supports a protocol called <a href=\"https:\/\/www.openarchives.org\/pmh\/\">OAI-PMH<\/a>. Luckily I had previously written a suite of software used to harvest all the bibliographic information and content from (OJS) OAI-PHM sites. Consequently, in a matter of about 30 minutes, I was able to create <a href=\"http:\/\/sites.nd.edu\/emorgan\/files\/2022\/04\/istl.csv\">a CSV file<\/a> listing all the articles along with their authors, titles, abstracts, URLs, etc.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I then ran a program that looped through the CSV file and downloaded (cached) the content. Thus, all the articles in their original form are found in the <a href=\"http:\/\/dh.crc.nd.edu\/tmp\/istl\/media\/istl.zip\">(temporarily) linked .zip file<\/a>. There are about 900 of them.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I then ran the whole thing through my <a href=\"https:\/\/reader-toolbox.readthedocs.io\">Distant Reader Toolbox<\/a>, and I am now able to characterize the journal as a whole. For example, after removing bogus files, there are about 850 articles, and the whole corpus is 2.5 million words long. (The Bible is about .8 million words long.) I was then able to create a <a href=\"http:\/\/dh.crc.nd.edu\/tmp\/istl\/media\/bibliography.txt\">rudimentary bibliography<\/a>, which is really only half a step better than the original CSV file.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Do you know what <cite>ISTL<\/cite> is about? Science and technology librarianship would be a good guess, but can you elaborate? I can, in a number of ways. For example, if I compute statistically significant keywords against the text, I can visualize the result as a word cloud. Now you know more, and in what proportions.<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img loading=\"lazy\" decoding=\"async\" src=\"http:\/\/sites.nd.edu\/emorgan\/files\/2022\/04\/keywords-cloud.png\" alt=\"keywords-cloud\" width=\"320\" height=\"240\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Rudimentary clustering of the data returns two possible themes, but the clustering process (Principle Component Analysis) does not articulate what those themes may be. Still, such an analysis points to what a good topic model might be.<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img loading=\"lazy\" decoding=\"async\" src=\"http:\/\/sites.nd.edu\/emorgan\/files\/2022\/04\/cluster-dendrogram.png\" alt=\"clusters\" width=\"320\" height=\"240\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Topic modeling with only two topics, returns two possible, over-arching themes: 1) students, and 2) search. Notice how the students theme is really about people, and the search theme seems to be about searching stuff:<\/p>\n\n\n\n<pre class=\"wp-block-preformatted\">       topic  weights                                           features\n    students  0.54019  students data research librarians faculty univ...\n      search  0.44664  search journals web access research articles d... \n<\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">In terms of proportions, the pie chart of the weights mirrors the clustering analysis.<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img loading=\"lazy\" decoding=\"async\" src=\"http:\/\/sites.nd.edu\/emorgan\/files\/2022\/04\/topic-model-01.png\" alt=\"topic-model\" width=\"320\" height=\"240\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">After a bit more modeling, the idea of search is still evident, but the theme of students has broken down into different types of people and different things being searched:<\/p>\n\n\n\n<pre class=\"wp-block-preformatted\">       topic  weights                                           features\n         use  0.34272  use faculty survey services new staff university \n      search  0.33438  search web database users results databases also \n        book  0.27892   book technology work new internet example many \n  librarians  0.25108  librarians research technology university educ...\n      access  0.20613  access journals electronic open research publi...\n    citation  0.19648  citation study journals research analysis arti...\n    students  0.17608  students literacy research instruction course ...\n        food  0.10850  food site resources environmental research agr...\n        data  0.10367  data research management researchers gis servi...\n        site  0.09101  site resources links web provides history rese...\n      patent  0.06577  patent yes patents titles databases journals d...\n   chemistry  0.05041  chemistry chemical structure molecular data bi...\n<\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Again, a pie chart of the whole.<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img loading=\"lazy\" decoding=\"async\" src=\"http:\/\/sites.nd.edu\/emorgan\/files\/2022\/04\/topic-model-02.png\" alt=\"topic model\" width=\"320\" height=\"240\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">One of our esteemed colleagues &#8212; Roy Tennant &#8212; once said, &#8220;<a href=\"https:\/\/rocky.uta.edu\/presentations\/handout.pdf\">Librarians like to search. Everybody else likes to find.<\/a>&#8221; Consequently, I was not surprised to see search as a theme in a library-related journal.<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img loading=\"lazy\" decoding=\"async\" src=\"http:\/\/sites.nd.edu\/emorgan\/files\/2022\/04\/quote.png\" alt=\"quote\" width=\"320\" height=\"240\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Upon closer inspection of the keywords, the word soil piqued my interest, so I created a full text index and searched for &#8220;title:soil OR keyword:soil&#8221;. I got three records:<\/p>\n\n\n\n<pre class=\"wp-block-preformatted\">  Your search (title:soil OR keyword:soil) against the study carrel named\n  \"istl\" returned 3 record(s):\n\n           id: 2495\n       author: Pellack, Lorraine J.\n        title: Soil Surveys \u2014 They're Not Just for Farmers.\n         date: 2009-09-01\n      summary: Soil surveys do contain inventories of the soils of an area;\n  however, they also contain a wealth of tabular data that help interpret\n  whether a location is suitable for a given use, such as a playground, a\n  golf course, or a highway. This guide will describe soil surveys, their\n  uses, and uniqueness.\n   keyword(s): library; soil; surveys; u.s\n        words: 3003\n     sentence: 151\n       flesch: 59\n        cache: \/Users\/eric\/Documents\/reader-library\/istl\/cache\/2495.htm\n          txt: \/Users\/eric\/Documents\/reader-library\/istl\/txt\/2495.txt\n\n           id: 2420\n       author: Bracke, Marianne Stowell\n        title: Agronomy: Selected Resources\n         date: 2007-07-01\n      summary: This web bibliography, or webliography, contains links and\n  descriptions of agronomy web sites that cover general and background\n  information, crop science, soil science, resources for K-12 teachers,\n  databases, freely-available and subscription-based journals, and\n  organizations. Only a select number of sites that focused on crop\n  science, soil science, or a particular sub-area (e.g., corn) were\n  included due to the large number of sites in existence.\n   keyword(s): agronomy; crop; information; plant; science; site; soil\n        words: 5135\n     sentence: 265\n       flesch: 39\n        cache: \/Users\/eric\/Documents\/reader-library\/istl\/cache\/2420.htm\n          txt: \/Users\/eric\/Documents\/reader-library\/istl\/txt\/2420.txt\n\n           id: 1984\n       author: Harnly, Caroline D.\n        title: Sustainable Agriculture and Sustainable Forestry: A\n  Bibliographic Essay: Theme: All Topics\n         date: 2004-08-14\n      summary: The authors found that there is no clear preference in\n  the marketplace to the many approaches to achieving sustainable forest\n  management. Peter F. Ffolliott, et al.'s book, Dryland Forestry,\n  details how to manage both the biophysical and socioeconomic aspects\n  of environmentally sound, sustainable forest management in dryland\n  environments.\n   keyword(s): agricultural; book; edited; farming; food; forest; forest\n  management; management; new; papers; press; soil; sustainability;\n  sustainable; sustainable agriculture; sustainable forestry; systems; topics\n        words: 14403\n     sentence: 990\n       flesch: 43\n        cache: \/Users\/eric\/Documents\/reader-library\/istl\/cache\/1984.htm\n          txt: \/Users\/eric\/Documents\/reader-library\/istl\/txt\/1984.txt\n<\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Well, that&#8217;s enough for now, but the point is this:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\"><p>As librarians we collect, organize, preserve, and disseminate data, information, and knowledge. These are the whats of librarianship, and they change very slowly. On the other hand, the hows of librarianship &#8212; card catalogs versus OPAC, MARC versus linked data, licensing versus purchasing, reference desking versus zooming, just-in-time collection versus just-in-case collection, etc. &#8212; change much faster with changes in the political environment and technology. Harvesting things from the Web and adding value to the resulting collection may be things we ought to do more actively. The things outlined above are possible examples.<\/p><\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">Fun with librarianship?<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">P.S. Ironically, I just noticed the <a href=\"https:\/\/blog.dshr.org\/2022\/03\/dangerous-complacency.html\">linked blog posting about Web archiving<\/a>. From the concluding paragraph:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\"><p>There is clearly value in identifying the parts of &#8220;the Web&#8221; that aren&#8217;t being collected, preserved, and disseminated to scholars. But in an era when real resources are limited, and likely shrinking, proposals to address these deficiencies need to be realistic about what can be achieved with the available resources. They should be specific about what current tasks should be eliminated to free up resources for these additional efforts, or the sources of (sustainable, not one-off) additional funding for them.<\/p><\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">Food for thought.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>I asked a colleague (Parker Ladwig) if there was a blog he thought might be worthy of archiving, and he mentioned Issues in Science and Technology Librarianship (ISTL). I took it upon myself to see what I could do. After looking more closely at the site, I guessed the underlying technology was not blog technology [&hellip;]<\/p>\n","protected":false},"author":92,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[10716],"tags":[],"class_list":["post-1052","post","type-post","status-publish","format-standard","hentry","category-distant-reader"],"_links":{"self":[{"href":"https:\/\/sites.nd.edu\/emorgan\/wp-json\/wp\/v2\/posts\/1052","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/sites.nd.edu\/emorgan\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/sites.nd.edu\/emorgan\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/sites.nd.edu\/emorgan\/wp-json\/wp\/v2\/users\/92"}],"replies":[{"embeddable":true,"href":"https:\/\/sites.nd.edu\/emorgan\/wp-json\/wp\/v2\/comments?post=1052"}],"version-history":[{"count":7,"href":"https:\/\/sites.nd.edu\/emorgan\/wp-json\/wp\/v2\/posts\/1052\/revisions"}],"predecessor-version":[{"id":1066,"href":"https:\/\/sites.nd.edu\/emorgan\/wp-json\/wp\/v2\/posts\/1052\/revisions\/1066"}],"wp:attachment":[{"href":"https:\/\/sites.nd.edu\/emorgan\/wp-json\/wp\/v2\/media?parent=1052"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/sites.nd.edu\/emorgan\/wp-json\/wp\/v2\/categories?post=1052"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/sites.nd.edu\/emorgan\/wp-json\/wp\/v2\/tags?post=1052"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}