Saturday, 21 August 2010

The psychological disorder of the filing system

From the title poem of George Szirtes excellent 2009 collection The Burning of the Books:

Librarian of the universal library, have you explored
The shelves in the stockroom where the snipers are sitting,
The repository of landmines in the parking bay,
The suspicious white powder at the check-out desk,
The mysterious rays bombarding you by the photocopier,
The pyschological disorder of the filing system
That governs the paranoid republic of print
In the wastes of the world?

Hoping that the "universal library" is not a veiled reference to the UL (we hold a collection of Szirtes' letters so I imagine he feels kindly towards us), library "filing systems" are often considered confusing, if not "psychologically disordered", by library users.

The Wikipedia entry tells us that "Classification systems in libraries generally play two roles. Firstly they facilitate subject access by allowing the user to find out what works or documents the library has on a certain subject. Secondly, they provide a known location for the information source to be located (e.g. where it is shelved)".

Perhaps these dual roles are at the heart of the confusion. Unsurprisingly they often prove to be incompatible - subject groupings being countermanded by physical factors like size and space. And for many works, subject classification decisions can seem arbitrary to a library (or indeed bookshop) user. In any case, the majority of bibliographic records contain headings which are able to record subject information with far greater complexity than a single call number.

I have argued elsewhere that users approach our catalogues knowing exactly what they want. If discovery isn't something which normally happens when browsing, then all a classification scheme needs to do is assign a fairly unique identifier to each item and provide a map showing the physical organisation of these identifiers. This is already the case for closed-access collections where browsing is not a factor.

What about classification for online resources, which don't have a physical location? Many libraries either assign a generic call number for electronic material, or don't assign one at all. Discovery is handled mainly through search, and any subject-driven browse facility is generated by the subject headings within the record itself.

Which is all very well in the catalogue, but what if we need to group online resources for other reasons? One of the things we're looking into is pushing subject-relevant content to students. In order to do this we need to "classify" these resources by subject. Some options are:
  • Use existing subject headings in record - BUT these can vary enormously in terms of "width" and "depth" (i.e. how many and how specific they are). They are often assigned by libraries with the nature of their collections in mind - if you only have 10 physics books, "physics" is fine, if you have 10,000 you'll need to be more specific. Can we be sure that subject headings have been applied in a consistent manner across our online resources?
  • Tap into existing forms of recommendation such as reading lists/citations and reflect them in the catalogue - seems the ideal solution BUT difficult to get your hands on the data and link it through to bibliographic records - an "expensive" option.
  • Use circulation data to make recommendations (i.e. "other people on your course borrowed this, other people who borrowed this borrowed that") BUT difficult to track for online material which is not "borrowed" as such - can we easily track access in the same way?
There are wider questions - does this kind of recommendation produce a concentration on core texts at the expense of wider exploration of the collection? Or, handled correctly, might it lead people into areas they would not have otherwise considered?

There are plenty of questions, and I for one don't have any real answers. Perhaps if I hang around by the photocopiers the "mysterious rays" might spark some inspiration ...

Saturday, 31 July 2010

Blogging as "autosave for our entire culture"



Interesting talk by Scott Rosenberg (one of the Salon pioneers), who has written a useful history of blogging.

Thursday, 29 July 2010

One way street to the iPad, paywalls and linked data

Writing in 1925, Walter Benjamin senses a radical change, not just in the physical forms which contain writing, but also in the nature of writing itself:

"Just as this time is the antithesis of the Renaissance in general, it contrasts in particular to the situation in which the art of printing was discovered. For whether by coincidence or not, its appearance in Germany came at a time when the book in the most eminent sense of the word, the book of books, had through Luther's translation become the people's property. Now everything indicates that the book in this traditional form is nearing its end."

This from a section entitled "Attested Auditor of Books" in Benjamin's brilliant collection One-Way Street, written in his usual aphoristic (blogging?) style (from which I will, with apologies, quote extensively). The internet has also been credited with allowing cultural output to become "the people's property", and leading to important changes in the forms which that output takes.

Benjamin isn't talking about the internet, but he is talking about a change in the form of "print", driven and shaped by economic and technological forces:

"Printing, having found in the book a refuge in which to lead an autonomous existence, is pitilessly dragged out onto the streets by advertisements and subjected to the brutal heteronomies of economic chaos."

So there are further parallels. The production of text, once controlled by publishers (in the broadest sense), is now subject to different forces - the kind of economic chaos Rupert Murdoch is trying to tame with his paywall? Perhaps.

There follows a lovely passage about the perpendicularity of text:

"If centuries ago it began to gradually lie down, passing from the upright inscription to the manuscript resting on sloping desk before finally taking to bed in the printed book, it now begins just as slowly to rise from the ground. The newspaper is read more in the vertical than the horizontal plane, while film and advertisement force the printed word entirely into the dictatorial perpendicular."

There's no argument that electronic content has, in the past, been mainly consumed on upright monitors in the "dictatorial perpendicular". There has been a lot of toing-and-froing over how effectively e-readers and the like can mimic "real" books (screen brightness, electronic ink) - I wonder how much thought has been given to the angle of reading and how it affects our consumption of print. Do the Kindle and the iPad herald an era when texts will once again cosily recline on their beds?

(Lacking either an iPad or a Kindle, I attempted to mimic "horizontal reading" by laying my monitor flat on the desk - don't try this at home. And yes, it does seem to make an immediate difference to one's attitude to the text, at least for this reader.)

If all this seems a bit prophetic - how about this for child/internet anxiety, 1925-style?

"... before a child of our time finds his way clear to opening a book, his eyes have been exposed to such a blizzard of changing, colourful, conflicting letters that the chances of his penetrating the archaic stillness of the book are slight"

It might sound like Benjamin is fondly harking back to this "archaic stillness" - but read on:

"... the book is already, as the present mode of scholarly production demonstrates, an outdated mediation between two different filing systems. For everything that matters is to be found in the card box of the researcher who wrote it, and the scholar studying it assimilates it into his own card index."

Not content with scholarly research databases, Benjamin goes straight onto linked data:

"It is quite beyond doubt that the development of writing will not infinitely be bound by the claims to power of a chaotic academic and commercial activity; rather, quantity is approaching the moment of a qualitative leap when writing ... will take sudden possession of an adequate factual content."

So words will not just carry the "ordinary" meaning that they have in text, but will be imbued with further meaning by the nature of their representation. Anyone who has worked with XML, let alone RDF (a "method for the conceptual description and modelling of information") will find this kind of thing familiar.

Signing off, Benjamin says that "poets" will be the new masters of language, implying that an understanding of the deep and various meaning of words will once again become of prime importance in a new system of communication. Perhaps we all need to be poets in the modern world. Again, his words seem prophetic:

"With the foundation of an international moving script they [poets] will renew their authority in the life of peoples, and find a role awaiting them in comparison to which all the innovative aspirations of rhetoric will reveal themselves as antiquated daydreams."

In the jargon of the poetry workshop, or equally of the linked data practitioner "don't tell - show!"

Anyone in search of a real treat could do worse than track this book down - if only for the next section, subtitled "Principles of the Weighty Tome, or How to write Fat Books" - point II as an example:

"Terms are to be included for conceptions that, except in this definition, appear nowhere in the book"

Or the wonderful, and thankfully, online Writer's Technique in Thirteen Theses

"Consider no work perfect over which you have not once sat from evening to broad daylight"

So much for Benjamin the blogger!

Wednesday, 7 July 2010

Context and meaning in search

My 18 month old son has just discovered sentences - or, more specifically one sentence - "Where's it gone?" (pronounced as one word - "Wezzigonn?" ). He throws his ball into the bushes. He looks at me dolefully and says "Wezzigonn?". I retrieve the ball. He throws it into the bushes. He looks at me dolefully and says ... well, you get the picture. Essentially, he thinks "Where's it gone?" is a single word which means "Fetch the ball, Dad". And the interesting thing is that in the context of the "game" I understand exactly what he means.

In Wittgenstein's Philosophical Investigations he posits a "language game" involving a builder and his assistant. Every time the builder needs another slab he shouts "Slab!" and his assistant duly brings him a slab. So the single word "Slab!" functions as the sentence "Bring me a slab!" in this context.

(He also discusses a variation where every time the builder needs a slab he calls out "Bring me a slab!". An observer who doesn't speak the same language assumes that "Bring me a slab!" is the word for "slab" and when building a wall himself calls to his assistant "Pass me a 'bring-me-a-slab'!" These kind of misunderstandings are commonly found in placenames - such as Bredon Hill, meaning "Hill Hill Hill".)

I suppose the point of all this is that context gives language meaning. Which is all very well if you're building a house, buying some cabbages, throwing a party etc. But what about when you "speak" into a search box. Where do you get your context from then? Do you have to play around with clever search modifiers so the interface understands that when you search Google Images for "bondage" you're looking for pictures of serfs?

Not entirely - at least not in Google and not in Amazon, which both use contextual information to give you relevant results. [In]famously, these interfaces give very different results depending on whether you are logged in. People get pretty uncomfortable with the whole idea of this 'contextual information' - how it's gathered, where it's stored. But the use of contextual information to give language meaning is an essential part of communication.

Libraries have thousands of users and millions of resources crying out to be introduced to each other. But our search mechanisms tend to be context-less. Lorcan Dempsey has said Discovery happens Elsewhere. For users of university libraries, discovery happens in context-heavy environments such as reading lists, citations, seminars, lectures. By the time they get to our interfaces they know exactly what they want. Then they find it (or don't).

Can we start to build some context into our own systems? And what kind of context would be useful? We can say straight off that for students the most useful context is what course they're doing (something we will soon have access to, and which I've blogged about elsewhere). If we also have access to course materials (i.e. reading lists) we can really start to provide useful context for searches. How about if we have access to the content of books and articles - in particular the citations they contain? Could we start to put our searches in the context of a scholarly network based on citation?

Or do we run the risk of second-guessing what users are searching for, and getting it wrong? There are endless anecdotes about people changing their relationship status on Facebook and immediately being bombarded with ads for wedding planners/speed dating. If people change course do we start serving up different results? And if discovery happens in context, how far should the library go in providing context, and how much should it leave to others?

PS Emma has pointed out that little Henry is playing the Fort-Da game http://www.cas.buffalo.edu/classes/eng/willbern/BestSellers/Catcher/FortDa.htm and that along with Lorcan and Wittgenstein, Freud could also be added to the list of tags!

Tuesday, 29 June 2010

Library Futures - A Surfeit of Future Scenarios

Earlier today, I came across an ACRLog post entitled Add Cyberwar Contingencies To Your Disaster Plan that includes a couple of links to reports on possible futures/scenarios that can help in planning future library needs and services:

- Futures Thinking for Academic Librarians, an announcement post for an ACRL report on “Futures Thinking for Academic Librarians: Higher Education in 2025”
- 2010 top ten trends in academic libraries, a "review of the current literature" by the ACRL Research, Planning and Review Committee.

These in turn reminded me of a couple of scenarios I'd heard that had been developed for JISC et al's Libraries of the Future project, and a quick dig around turned them up here: Libraries of the Future - Outline Scenarios and Backcasting

The Libraries of the Future project has identified three scenarios:

- Wild West, introduced as follows:
2050 is an era of instability. Governments and international organisations devote much of their time to environmental issues, aging populations and security of food and energy, although technology alleviates some of the problems by allowing ad hoc arrangements to handle resource shortages and trade. In this environment, some international alliances prosper but many are short term and tactical. The state no longer has the resources to tackle inequality, and is, in many cases, subservient to the power of international corporations and private enterprise.

The challenges of the 21st century have created major disruptions to academic institutions and institutional life. Much that we see as the role of the state in HE today has been taken over by the market and by new organisations and social enterprises, many of them regional.

- Beehive, introduced as follows:
The need for the old European Union countries to maintain their position in the world and their standard of living in the face of extensive competition from Brazil, Russia, India and China (BRIC) has led to the creation of the European Federation (EF) under the treaty of Madrid in 2035. The strength of the EF has meant that values in the EF have remained open in the long tradition of western democracy and culture.

In the years leading up to 2050 the world became increasingly competitive; the continuing economic progress of the BRIC countries and their commitment to developing high quality HE systems means that even high-tech jobs are now moving from the West. On a worldwide scale, and in the US, UK and Europe especially, employer expectations now dictate that virtually all skilled or professional employment requires at least some post-18 education. In the UK these drivers have resulted in a state-sponsored system that retains elements of the traditional university experience for a select few institutions while the majority of young people enter a system where courses are so tightly focused on employability they are near-vocational.

- Walled Garden, introduced as follows:
Following the global recession of the early 21st century cuts in investment levels to help reduce the national deficit meant that internationally, the UK’s influence waned and it became ever more isolated. Indeed the UK drifted from the EU, particularly after the Euro collapsed in the century’s second global recession, and the UK itself fragmented as continued devolution turned to separation and independence. Fortunately, the home nations have achieved reasonable self-sufficiency.

Technological advances, whilst allowing some of the challenges faced earlier in the century to be overcome, has also brought its
problems. The ability for people to connect with like-minded individuals around the world has led to an entrenchment of firmly held beliefs, closed values and the loss of the sense of universal knowledge. This has resulted in a highly fragmented HE system, with a variety of funders, regulators, business models and organisations that are driven by their specific values and market specialisation. However, ‘grand challenges’ of national importance goes some way to galvanising the sector.


The ACRL 20205 report identifies 26 possible scenarios (?! - I thought the idea of scenario planning was to identify a few that covered the bases between them?!), with a range of probabilities of them occurring, their likely impact, and their "speed of unfolding" (immediate change, short term (1-3 years), medium term (3-10 years), long term (10-20 years)).

High impact, high probability scenarios include:

- Increasing threat of cyberwar, cybercrime, and cyberterrorism, introduced as:
College/university and library IT systems are the targets of hackers, criminals, and rogue states, disrupting operations for days and weeks at a time. Campus IT professionals seek to protect student records/financial data while at the same time divulging personal viewing habits in compliance with new government regulations. Librarians struggle to maintain patron privacy and face increasing scrutiny and criticism as they seek to preserve online intellectual freedom in this climate.

- Meet the new freshman class, introduced as:
With laptops in their hands since the age of 18-months old, students who are privileged socially and economically are completely fluent in digital media. For many others, the digital divide, parental unemployment, and the disruption of moving about during the foreclosure crisis of their formative years, means they never became tech savvy. “Remedial” computer and information literacy classes are now de rigueur.

- Scholarship stultifies, introduced as:
The systems that reward faculty members continue to favor conventionally published research. At the same time, standard dissemination channels – especially the university press – implode. While many academic libraries actively host and support online journals, monographs, and other digital scholarly products, their stature is not great; collegial culture continues to value tradition over anything perceived as risky.

- This class brought to you by…, introduced as:
At for profit institutions, education is disaggregated and very competitive. Students no longer graduate
from one school, but pick and choose like at a progressive dinner party. Schools increasingly specialize by offering online courses that cater to particular professional groups. Certificate courses explode and are sponsored by vendors of products to particular professions.


The 2010 top trends from the literature review are given in no priotised order as:

- Academic library collection growth is driven by patron demand and will include new resource types.
- Budget challenges will continue and libraries will evolve as a result.
- Changes in higher education will require that librarians possess diverse skill sets.
- Demands for accountability and assessment will increase.
- Digitization of unique library collections will increase and require a larger share of resources.
- Explosive growth of mobile devices and applications will drive new services.
- Increased collaboration will expand the role of the library within the institution and beyond.
- Libraries will continue to lead efforts to develop scholarly communication and intellectual property services.
- Technology will continue to change services and required skills.
- The definition of the library will change as physical space is repurposed and virtual space expands.

What strikes me about all these possible scenarios is that there don't seem to be any helpful tools that let you easily identify and track indicators relating to the emergence of particular aspects of the scenarios, which I think is the last step in the process of scenario development espoused in Peter Schwartz's "The Art of the Long View"?

So for example, OCLC recently released a report called A Slice of Research Life: Information Support for Research in the United States, which reports on a series of interviews with research and research related staff on "how they use information in the course of their research, what tools and services are most critical and beneficial to them, where they continue to experience unmet needs, and how they prioritize use of their limited time." And towards the end of last year, the RIN published a report on Patterns of information use and exchange: case studies of researchers in the life sciences (a report on information use by researchers in the humanities is due out later this year(?), and one for the physical sciences next year(?)...) A report on researchers' use of "web 2.0" tools is also due out any time now...

So, are any of the trends/indicators that play a role in the 2025 scenarios (which are way too fine grained to be useful?) signaled by typical responses in the OCLC interviews or the Research Information Network report(s)?

PS As if all that's not enough, it seems there's a book out too - Imagine Your Library's Future: Scenario Planning for Information Organizations by Steve O'Connor and Peter Sidorko. (If the publishers would like to send me a copy...?! Heh heh ;-)

Friday, 25 June 2010

I've Got Google, Why Do I Need You?

An excellent presentation on how a modern student percieves the way a library works. Its a great reminder on the gap between the web native student experience and the traditional library service. It also has my favourite quote of the week: "The librarian's logic is just as alien to me as the programmers' logic" ...

By way of Angela Fitzpatricks' shiny new blog!

Saturday, 19 June 2010

a wealth of reference management

There's a lot going on in the field of reference management tools - especially here in Cambridge.

Reference management tools include all kinds of systems which help you organise references you have found, store the papers they refer to and perhaps annotate them, share the citations with others, cite papers in your own works, and so on.  There are some big players in this area - the first two which spring to mind for me are Zotero and Mendeley. Zotero is a Firefox plugin, so it sits within your browsing experience; Mendeley is a website and a downloadable tool, and makes a big effort to connect you to others and recommend other works - the "Last.fm of scholarly work". Both are popular with researchers at the University of Cambridge.

I realised this week that I now know of at least four reference management tools just originating here in Cambridge:

  • Papers. This is Mac software from Mekentosj, and has recently won an award
  • iCite. Like Zotero, this is another Firefox plugin
  • PaperPile, an open source system (GPL) from the EBI, which was first written for Linux
  • qiqqa, a somewhat unpronouncable name for a Windows application. 
It's great to see the local entrepreneurial spirit coming into play in the academic sphere!
    These are a tiny fraction of the world of reference management. It's interesting to note that the market can support so many tools. Each has special features which will appeal more to some users than others; some are particularly well suited to one discipline, with better support for their paper types and bibliographic databases. Of course, it's possible to combine two or more tools as part of your scholarly workflows, to get the best bits of each...