Friday, 30 October 2009

What Is the Role of the Library in an "Open" University?

A post on the freeculture blog today (Call for Participation: Join the Open University Campaign!) raises a call to arms "to increase collaboration, sharing, and openness at the level of higher education" and proposes a report card that can be used to assess the level of openness of a Higher Ed institution.

The reporting is proposed around the five tenets of the Wheeler declaration(?):
"An open university is one in which:
  1. The research produced is open access;

  2. The course materials are open educational resources;

  3. The university embraces free software and open standards;

  4. The university’s patents are readily licensed for free software, essential medicine, and the public good;

  5. The university’s network reflects the open nature of the Internet,

where “university” includes all parts of the community: students, faculty and administration."


Looking at these aims, I wonder to what extent they might relate to the role of the academic library in a digital world? So for example, is there a role for the Library as the champion of - and gateway to - openness (as described above)? Or should openness initiatives be the remit of other parts of the institution?

Just by the by, author and activist Cory Doctorow is in Cambridge next week, speaking at a free and open Arcadia Seminar on Tuesday, 3rd November in the Umney Theatre, Robinson College, Cambridge. For more details, see Arcadia Seminar (3rd Nov): "Thinking Like a Dandelion: Cory Doctorow on copyright, Creative Commons and creativity".

Thursday, 29 October 2009

e-Books and Content Packaging

One thing I've noticed wandering around the bookshops of Cambridge is that they have started putting up Point-of-Sale marketing displays pushing e-book readers. I'm still not sure whether these will fly or not, though I can see the form factor is more accommodating than reading ebook content on a mobile phone or iPod touch, because I'm not sure what the future of the book is...

For sure, I'm a fan of books (our house is full of them, I buy them regularly, new and old, I borrow them from libraries (note to self: renew overdue library books) and so on), but I'm not convinced that ebooks will necessarily work like books when it comes to buying them. Just like I'm not sure that e-journals are necessarily like print journals when it comes to buying them. Becuase in an academic or technical context at least, books and journals aren't so much as read, as dipped into...

So for example, textbooks and reference manuals aren't necessarily read cover to cover - they tend to be dipped into at a topic, or problem level.

And journals aren't typically read cover to cover, nor are the papers within them necessarily read from start to end. Taking a leaf out of Peter Murray Rust's book, I've been chatting to folk over the last week or so and asking them how they read journal papers - a skim of the abstract, a quick look over the conclusions, and possibly a deep dive into a particular section seems to be as common a way as any.

So with a library purchasing policy based on buying (and lending) whole books, and subscribing to whole journal sets (not just single journals), how does this compare with the way the information contained within these content packages is actually referred to? Not very well, I think?

The origins of this idle thought lay in part with a TechCrunch splash earlier today about DeepDyve, also reported in DeepDyve — iTunes comes to Science Publishing:



("Article renting" - how does that sound to you?!;-) I haven't looked at the detail of this rental service yet, but with competitive pricing ("For just $0.99 per article, users of this 'pay-as-you-go' plan can rent and read a premium article from one of the many prestigious journals available through DeepDyve. Articles can be read multiple times for up to 24 hours.") I think I can see what they're trying to do... (which is to make the fee bearable for the convenience of accessing the content that is being sold... (that is, the convenience that is being sold...).)

Remembering the various debates around whether services like Google Scholar (and the ill fated Microsoft Live Academic search*) offered a comparable service to subscription databases in terms of the content that was being indexed, I wonder how well the list of publishers/publications that DeepDyve is indexing will stand up?

(* Hmmm.... seems that Microsoft has a research search engine for computer science literature? Microsoft Academic Search)

The TechCrunch post was presumably a lazy response (I hold TechCrunch in very low regard...) to a PR push (and this press release) from DeepDyve following the announcement of Safari Books Online 6.0, which offers subscription access into O'Reilly's online technical books content. (Disclaimer - I buy loads of O'Reilly books... And on the convenience ticket again, I know several members of HEIs who have taken out personal subscriptions to Safari to get round the limited access subscriptions that many institutions have to Safari.)

Looking round the O'Reilly site a bit further, their e-book offerings try to cover the range of likely candidate formats - mobi for the Kindle, PDF for laptops, epub for mobile devices.



In a tweet earlier today, @ostephens also pointed out there is a keen pricing policy on O'Reilly titles as iPhone apps



And as if to close the circle, there is also an appstore on the O'Reilly domain - O'Reilly Best iPhone Apps:



But what's this - no books category???

So where are we at? When it comes to ebooks, don't think of them as books... to get your creative juices flowing, it may be more rewarding to think of them as apps... ;-)

Wednesday, 28 October 2009

Looking Out for "Linked Course Data"

In The Library's Role in Organising "Course Knowledge", I suggested that one possible role for the library might be in organising the knowledge an institution has, or accretes, around the institution's courses over a period of time.

This knowledge might include things like reading lists, and past exam papers, for example.

Looking round the .cam.ac.uk website, it strikes me that the Computing Laboratory have already collected a lot of relevant material around their teaching, although not necessarily in a portable format.

So for example, we have a breakdown of syllabus information for each lecture course associated with the Computer Science Tripos for the current academic year:



Each topic has a page associated with it outlining the syllabus, as well as learning objectives and recommended reading:



The format of the URIs is based on a strightforward enumeration: http://www.cl.cam.ac.uk/teaching/0910/CST/node13.html in the above case, for example.

A second area of the site provides a slightly different breakdown of the teaching materials associated with each lecture course:



In many cases, it seems as if the URI is often 'inspired' by the title of the course, e.g. http://www.cl.cam.ac.uk/teaching/0910/DiscMathI/ in the above case.

Included on the course materials page is information that identifies which other programmes (that is, "Subjects" in the CAMSIS parlance) aside from the Computer Science Tripos expect students to take papers/exams in these subjects.

Note that whilst the Subject codes are not explicitly identified on either the syllabus page or the course materials page, they are hinted at in the URI of the syllabus pages, and in the Taken By area of the course materials page. In the above we see CST for example, which is close to the CAMSIS subject code:



Back to the course materials page, and the links to past papers also tend to include a human readable element: http://www.cl.cam.ac.uk/teaching/exams/pastpapers/t-DiscreteMathematicsI.html although it is not necessarily the case that you can directly transform the URI of the past papers page from the course materials page.

Note that there doesn't seem to be a way of identifying the URI of the syllabus page from the course material or past papers page (or vice versa).

If we look at the past paper areas:


and in more detail at an actual past paper:



there is only 'anecdotal' evidence of the CAMSIS examination code (i.e. 'Paper I').



The URI of the past papers, whilst readable to an extent, are also not forthcoming in this respect: e.g. http://www.cl.cam.ac.uk/teaching/exams/pastpapers/y2009p1q4.pdf

So where are we at in terms of being able to use URIs and CAMSIS codes as "Linked Data"? (see also: CAMSIS Codes "-ish Linked Data" Map)




Nowhere we'd want to start from, I suspect...;-)

So what would have been nice? Well what have we got?

- Subject codes (for the Computing Science Tripos, Natural Sciences Tripos, etc), that look something like CST; (there are also UCAS codes, which I suspect must map on to Subject codes, somehow?)
- Parts (roman numerals, maybe with an A/B) (which sort of map on to year of study... sort of...); note that subject code + part already have CAMSIS codes defined for them;
- Papers (a number), which relate to exams;
- (Lecture) courses (titles, numreous incosistent short forms), which map onto a particular paper in the context of a Part and a Subject.

So what would be nice would be:
- a unique identifier for each lecture course;
- a mapping from a lecture course identifier to past paper, keyed by Subject, part and exam number. Note that (like a DOI) this might be a many to one mapping. So for example, Discrete Mathematics I, which is taken by: Part IA CST, Part IA NST, Part I PPS, might have require four plus one pieces of information in its identifier: e.g. 2009/CST0/1/DM_I for the year, subject and part, paper and course in the Computing Science Tripos, and 2009/CST0/1/DM_1:NST0 for its NatSci variant?

The question number in a particular exam paper might also be handy in cases where each different lecture course has it's own identifiable subsection within a paper.

In other institutions, it is likely that the course (module?) will have its own unique identifier. So for example, if we look at the Reading List system at Manchester, we can see how courses there are identified by a unique code within the first year Computing degree:




If we tunnel into the Plymouth Reading list system, we find that whilst module codes are identified, there is no link back to the parent degree code:



The RDF version of the page does appear to have an informal (i.e. not by code number) reference back to the 'parent' degree programme, though:



All of which is to say: at the moment there is no obvious/easy way to map from a UCAS code to an internal degree code to a year of study to a particular lecture course to its reading list and/or past papers. And in the Cambridge context, there's no obvious way of going from a URI associated with content relating to any one of those levels to a URI for content relating to any other ;-) (I think!)

Tuesday, 27 October 2009

IT Skills or Information Skills?

The closing question to the panel from my session at ILI2009 a couple of weeks ago related to what skills we thought would be most desirable in the librarian of the future. My response was something along the lines of 'be able to write a good query', the thinking being that if a researcher or undergraduate goes to a subject librarian asking for help, it might be useful if they went away with a good query they could run repeatedly as a saved (persistent, alert raising) search, rather than just leaving with one or two particular references.

So this got me thinking: what are the differences between information skills and information technology skills?

As far as I know, public commenting still doesn't work on this blog, so what skills are you going to have to invoke to be able comment on this post, and what skills am I going to have to use to find those comments...?

[I think public comments are now working???]

Peter Murray Rust at Internet Librarian 2009

http://vimeo.com/7130694

As thought provoking as ever, Peter outlines how he believes the current situation with publishers has ensures that Librarians no longer fulfil Ranganathan's five laws.

He also roughly describes how he would like science librarians to work, effectively as information brokers doing all the work scientists currently do by engaging with online communities, publishing in an open fashion and writing new tools.

The Library's Role in Organising "Course Knowledge"

It seems to me that one of the key, live discussion issues faced by academic libraries at the moment in their 'service to undergraduates' role is the extent to which they integrate with the students' experience of a course (which increasingly means: in the VLE).

In the Open University, the central library's role was originally to support course development and academic research, as well as archiving the university's course materials. Undergraduate access to other academic libraries across the UK was also negotiated by the Library. The rise of e-collections, however, means that the OU Library can now start to provide information access to undergraduate students who access library services remotely.

In Cambridge, too, the central University Library's role, primarily as a research library, but also as a provider of services to College and Department libraries, looks as if the increasing online availability of resources might influence the services it provides directly to undergraduates, in addition to the services it provides through the other libraries.

But to what extent are the libraries looking inside their respective .ac.uk domains (both the public areas and the authenticated ones) and attempting to 'organise learning related course knowledge' - that is, those resources that grow up around a course over the its lifetime.

So for example, we have 'direct' course related outputs, such as course reading lists and past exam papers. Then there are the outputs of data-mining around courses, such as looking at what books students on a particular course borrowed in significant numbers. Where students do their own research on a particular topic, there are resources they have referenced in their work. These may become increasingly discoverable if students start to make use of social bookmarking or reference managing tools, posting course related social bookmarks, for example, as they do their research. For students who subscribe to RSS feeds, OPML feedrolls relating to a particular course might also be available. And as open educational resource initiatives keep testing the waters in terms of releasing course materials under an open license, there are likely to be increasing amounts of course materials (including reading lists, lecture podcasts, or videos etc etc) available relating to any particular course, or at least, course related subject area.

(One OU course I used to mark each year always has a research style question that required students to discover two or three resources to support their answer. I'd bookmark these each year in order to see what sorts of searches students might be doing for the assessment: a lot of the resources seemed to be among the top Google results for an 'obvious' keyword/keyphrase selection suggested by the question.)

So how might a Library go about organising 'course knowledge'?

One approach I have explored in the past are course related search engines. In one early demo, I mined courses on the OU's OpenLearn website for links to external websites, and used these as the basis for a course resources search engine. (I suspect the demo has since rotter: OpenLearn Dynamic Custom Search Engine). In other examples, I would search over the outgoing links from a particular webpage (e.g. 'Search Links On this Page' Revised Bookmarklet - again, this may have rotted...). One course I am currently running uses a Google Custom Search Engine populated by hand with resources linked to from the course materials, providing students with a way of searching over the (public) resources that the course refers to (think of this like a reading list limited search engine).

"Ah yes, all very useful", you may say. "But so what...?" So maybe Google is going to beat the libraries to it again if they don't start thinking in a weblike way. For example, in my feed reader today I learn that Google is now offering Contextual search within Wikipedia:
For Wikipedia pages with a lot of information and links, contextual search lets you limit your search to only those Wikipedia pages that are linked from the current article, focusing the results on the topic of the article. So, in addition to getting all matching Wikipedia articles, you can quickly drill down to contextually relevant results using the Linked Wikipedia Pages tab.

WAKE UP: in most institutions, the course is a context, but how much value is being exploited from that context? To what extent do our 'learning environments' allow linked course resources to be explored as if the learning environment was a research environment?

Encouragingly, there are signs of academic Libraries starting to think about the exploitation of course resources. So for example, @daveyp's book recommender based on library loans data looks as if it might be increasing the number of books borrowed at Huddersfield University (ILI2009: Exploiting Usage Data). And at Cambridge, moves appear to be afoot at getting hold of centralised course information data so that it can be used to personalise a variety of library services (more on this in a later post...;-)

Just by the by, clearing out some long ago opened tabs on my web browser ovr the weekend, I came across this post entitled Examopedia, which describes a project at Portsmouth where "questions from a past exam [are posted on a wiki]; students then put their answers to these questions onto the wiki, suggesting alternatives or amendments if an answer has already been posted".

For subjects where short problems are used as one of the drivers for checking understanding and demonstrating techniques, I can see how this might be a really powerful technique. It also got me wondering about the extent to which a service like Stack Overflow (now available as a white label service, Stack Exchange), where users can post questions, receive answers, and share all sorts of karmic goodness, might be used in a similar way? E.g. by allowing past papers to be atomised down to the level of separate questions, posted as such, and then collaboratively addressed by a social learning community?

One question is, to what extent should it be the Library's role to explore the way(s) in which course knowledge can best be organised and exploited...?

Thursday, 22 October 2009

JISC MOSAIC Competition Entries - Imaginings Around the Use of Library Loans Data

One of the joys of chatting to folk under the auspices of the Arcadia programme is that it allows them free reign to talk about all the things they'd like to do, rather than just the things that are on the current day-to-day grind development plan. Today, for example, I attended a meeting that was looking at ways of getting course related data associated with students held in one university administrative database so that it could be used to provide additional context for library services looking to make course related recommendations (reading lists, past exam papers, and so on) to individual students.

This relationship between books and courses is one that can be explored, in proof of concept prototype form at least, using data from the JISC MOSAIC programme.

Over the summer, this programme supported a developer competition to encourage developers to "produce a browser based application that makes use of some or all of the MOSAIC library activity data".

The data represented several years' worth of anonymised records relating to books borrowed from an academic library, with book loans linked to course affiliations of the borrowers, as well as their year of study.

Although the data was initially made available as a large XML file, Dave Pattern (@daveyp) from the University of Huddersfield Library rapidly produced a simple web based API to the database (MOSAIC Competition Data API) which opened up the possibility for developers and tinkerers without database experience to participate in the competition.

In all, six entries were submitted to the competition, (comptition results), covering the following areas:

Improving Resource Discovery

- Navigate the ‘Book Galaxy’ through links based on borrowing habits (link: Book Galaxy):

[image to come - I couldn't get the Java applet to work]

- "find books that have been borrowed from your institution's library by students on a particular course, add them to a reading list and share it with others" (link: iLib):

iLib MOSAIC Competition entry http://www.codebrane.com/ilib/course/bahonseducationaladministration

iLib glue: e.g. http://www.codebrane.com/ilib/book_search/?book_searchtext=games, http://www.codebrane.com/ilib/course_search/?course_searchtext=games

Supporting learning choices

- Supporting course choice - augmenting the UCAS course catalogue with books and courses related to a course by virtue of books borrowed commonly across courses(link) [DISCLAIMER: this was my entry; blog post about this app]:

Supporting course selection MOSAIC competition entry http://ouseful.open.ac.uk/mosaic.php?cc=e216

Course selection glue: e.g. http://ouseful.open.ac.uk/mosaic.php?cc=e216, and the various bits of Yahoo pipework associated with the app (listed on the project page)

- Course suggestions based on books you’ve read or that are itemised in a reading list (link: Read to Learn):

Read to Learn MOSAIC competition entry http://www.meanboyfriend.com/readtolearn/

Supporting decision making

- Assess circulation relating to departments and courses (link: Collection Development Dashboard):

Collection Dev Dashboard - MOSAIC Competition entry http://dysinterested.com/mosaicdashboard/


- Value the loans per courses as a collection performance indicator (link):

Book loan value MOSAIC competition entry http://voyager.aber.ac.uk/mosaic/

Aggregated Data
A couple of the apps display reports that crunch aggregated stats from the data in potentially useful ways (e.g. the collection reporting tool and the loan book value tool), although I'm not sure about the extent to which Library systems support bulk reporting already (library systems vendors all seem determined to keep their documentation behind password protected, license holder access only barricades for some reason that I just can't fathom?

The Read to Learn app looks like it combines lookups from several ISBNs to rank suggested courses, which to my mind is more interesting (in a playful, if not useful, sense) than simply totalling figures within a particular set. E.g. I assume that the suggested course rankings will be different for each different set of books you upload and are calculated uniquely for each different set of books? (If I've misinterpreted this, I'm sure Owen will let me know ;-)

"Linked Data" opportunities
On thing that particularly appealed to me about this data was the ability to recommend one thing based on another. In my own application, this took the form of finding courses on which books had been borrowed that were also borrowed on a course that a student browsing a course list on the UCAS website might be considering taking.

At first, I though iLib did something comparably indirect - e.g. using a search term to find a set of books and then using that list of books to identify a set of related courses, or searching for coursenames that contain a particular search term and then reporting back with the books associated with those courses; but I think the course search and book search are actually just literal searches? (i.e. the book search searches for books directly, and the course search searches for course titles containing the search term directly).

I'm not sure what Book Galaxy does from playing with it - I still can't get it to work :-( From the description: "Clicking on a book will show a web of related books for the selected book at the centre, with courses that use the book listed around the outside." So this presumably looks up books related to courses related to a particular book (nice:-). "Clicking on a course will show all the books that are used by that course." But not the courses related to that course by virtue of common books, which would be the corollary of the books related to books approach? (So e.g. my app used a courses to courses mapping.)

Just looking at first and second order relationships, we can get:

- course to books (starting with Course A):
Mosaic: course to book

- book to courses (starting with Book 1):
Mosaic - Book to courses

- course to courses (starting with Course A):
Mosaic: course to courses

- book to books (starting with Book 1):
Mosaic: Book to books

Note that if we take xISBNs into account (and the MOSAIC API supports xISBN lookups), we get a potentially richer map, e.g. when mapping from a book to courses. So for example a book has alternative ISBNs for its diffrent editions, we might get something like:

MOSAIC: Book to courses via xISBN


Anyway, it was interesting to see how different people addressed the competition, although there's not as much glue as I was hoping for (that is, apps with RESTful URIs that can take course codes, ISBNs or free text search terms, and maybe produce RSS, JSON, XML or easily scrapable outputs that don't simply replicate calls to Dave's API).

And as far as the data goes, I think MOSAIC is still accepting data for to add to the mix, and scripts are available (I think?) that can generate a data dump in the required format from a variety of library systems (including Voyager... So how about it, Cambridge?;-)

For more techie details, see the MOSAIC wiki

PS if any of the other competition entry developers have blogged about their apps, I'd love to be able to link to the corresponding posts :-)