Tuesday, 15 November 2011

ANCIL at LSE



Emma Coonan, Jane Secker and Helen Webster gave a great seminar presentation  at LSE on the new curriculum for Information Literacy that emerged from Emma's and Jane's Fellowships.  (Helen is now working on implementation issues.)

Tuesday, 25 October 2011

LOCKSS

LOCKSS (Lots of Copies Keep Stuff Safe), based at Stanford University Libraries, is an international community initiative that provides libraries with digital preservation tools and support so that they can easily and inexpensively collect and preserve their own copies of authorized e-content. LOCKSS, in its twelfth year, provides libraries with the open-source software and support to preserve today’s web-published materials for tomorrow’s readers while building their own collections and acquiring a copy of the assets they pay for, instead of simply leasing them. LOCKSS provides fully managed preservation and 100% post cancellation access.
LOCKSS site

Sunday, 2 October 2011

Ten questions for wannabee "digital scholars"

From Martin Weller's Blog:

Do you have an open access publishing policy for the outputs of the project?
Will you make the data openly available, or is some of it not appropriate to do so? If so, where will you put it?
Will there be one main individual in charge of social media communication or is it just distributed across the group?
Will you have specific twitter/blog/youtube accounts for the project or do individuals in the project have a high online reputation it would be better to utilise?
How will you incorporate analytics into the project? Do you expect to get a certain number of views, dwell time, global distribution, etc on a main site?
What media will you use? For example, will you create a 'trailer' video for the project and an overview of the findings?
How will you archive discussion around the research, eg twitter hashtag conversations?
Will you provide a curation service, eg a Scoop-It page of relevant resources as you go along?
Are there new methodologies you will employ, eg crowdsourcing?
Is there a planned release for findings throughout the project? Will any aspects be not open for dissemination eg via twitter or blogs?

Tuesday, 14 June 2011

John Palfrey's closing SSP Keynote

John Palfrey -- of Harvard Law School and the Berkman Center -- is taking a close interest in library issues and delivered the closing plenary session at this year’s SSP Annual Meeting. Kent Anderson has provided a useful summary of what he said.

Saturday, 4 June 2011

Hacking the Academy: contents list now available


The full list of contents for the edited volume is now available. Over 200 scholars took part in the project. All the contributions are online here.

Fascinating project.

Tuesday, 22 March 2011

Court rejects Google Books settlement

Significant setback in Google's path to world domination. CNET News reports that

Adding another chapter to a long, drawn-out legal saga, a New York federal district court has rejected the controversial settlement in a class-action suit brought against Google Books by the Authors Guild, a publishing industry trade group.


"While the digitization of books and the creation of a universal digital library would benefit many, the ASA would simply go too far," a court document explains. "It would permit this class action--which was brought against defendant Google Inc. to challenge its scanning of books and display of 'snippets' for on-line searching--to implement a forward-looking business arrangement that would grant Google significant rights to exploit entire books, without permission of the copyright owners. Indeed, the ASA (Amended Settle Agreement) would give Google a significant advantage over competitors, rewarding it for engaging in wholesale copying of copyrighted works without permission, while releasing claims well beyond those presented in the case."


The settlement would grant Google the right to display excerpts of out-of-print books, even if they are not in the public domain or authorized by publishers to appear in Google Books. When the settlement was initially announced in mid-2009, opposition flooded in from lawyers on behalf of Microsoft, the Electronic Frontier Foundation, and a coalition called the Open Book Alliance who decried it as anticompetitive.


"Google and the plaintiff publishers secretly negotiated for 29 months to produce a horizontal price fixing combination, effected and reinforced by a digital book distribution monopoly," a lawyer for the Open Book Alliance said at the time. "Their guile has cleared much of the field in digital book distribution, shielding Google from meaningful competition."

Well-known economics journal goes open access

Interesting post by Justin Wolfers:

Here’s a test of basic economic literacy: What is the socially optimal price of online access to economics journal articles?

If my students learn only one thing, it’s this: Price equals marginal cost. And the marginal cost of accessing a journal article is pretty much zero. The research has been written, the type has been set, and the salaries have already been paid — usually thanks to a university, think tank, or government grant. So the socially optimal price is: free. Every time we charge a price higher than this, we risk pricing out someone who might benefit from the insights of an academic scribbler.

The Brookings Papers on Economic Activity – the journal that David Romer and I edit — has decided to take this piece of economic wisdom seriously. The Brookings Papers are now entirely open access. Yep, we’re charging zero; nada; nothing; zip.

Sunday, 20 March 2011

The Internet-Informed Patient: March 28




We're holding a one-day symposium on March 28 on the implications of the Internet-informed patient for health care. The previous day (March 27) we will be hosting a Hack Day for developers and other professionals interested in harnessing IT to improve our use of health data -- both in terms of patient-centred Apps and making creative use of open health data to inform patients and improve decision-making.

Both events will be held in the Moller Centre at Churchill College, Cambridge.

The symposium will not take the usual form of a panel of expert speakers with a (largely passive) audience, but is designed as a day-long series of structured conversations in small groups primed by a few pithy, thought-provoking statements on a number of burning issues. The conversations will be live-blogged and some contributions will be video-recorded, so that a record of the conversations can be made available afterwards.

We still have some places available for both the Symposium and the Hack Day. If you're interested -- or know someone who might be -- please get in touch.

For the Symposium the contact is Isla Kuhn -- ilk21 (at) cam.ac.uk
For the Hack Day the contact is John Naughton -- jjn1 (at) cam.ac.uk

You can also get in touch via the website -- http://www.iip-symposium.info/

Friday, 7 January 2011

'Lending' Kindle books

The Kindle was apparently Amazon's best-selling product at Christmas, but many of its new users feel uncomfortable with the DRM-lockdown that comes with it. According to this report, though, there's been a slight relaxation -- in that you can 'lend' a Kindle book to a friend for 14 days.

Amazon is allowing Kindle users to lend a book to a mate, but the UK Publishers Association reckons e-book borrowers should get down the library.

The new feature allows e-books bought for the Kindle platform to be lent out for 14 days, delivered by email and springing back to their owners automatically as detailed by Amazon, but the Publishers Association (PA) is unlikely to approve, given its stance that anyone wanting to borrow an e-book from the local library should get their bones down to the building for a bit of physical interaction with their local community.

Friday, 17 December 2010

End of Delicious?

Many reports are circulating on the Interweb regarding Yahoo's planned closure of several sites, not least of which is Delicious, the social bookmarking site. From being one of the first web 2.0 successes, the site has had many problems over the past few years and some have noted its failure to adapt and evolve to meet changing expectations.

We used Delicious to list useful web resources on the first ever Arcadia project, science@cambridge. Many libraries in Cambridge and beyond have also done the same, its a great tool. Since then, the potential risk of loosing third party infrastructure like this has often popped up in discussion. Now it may be a reality. (Large portions of our site are also using Pipes. Lets also keep our fingers crossed for that superb service).

Thinking on a wider scale, Delicious, like Wikipedia, StackOverflow and many other online resources full fill some of the functions of a library in the networked world, namely the classification of units of online information. Many people rely on it daily, and much noise has made of its community basis as a real alternative to traditional means of classification.

Now thanks, to a corporate reshuffle, it may just disappear as a result of market conditions. I'm left with on a Friday afternoon with three things to think about:

  • Why was the site judged a failure? Is tagging a fad that will fade, whilst traditional classification will somehow endure (this I doubt) ? Is it because its function was better provided by other successor sites, or some other reason?
  • If the market cannot sustain these networked library-like services, should libraries (or the non-profit educational sector) start developing services like Delicious? Would we be better placed to provide this vital web infrastructure over a commercial entity? Would it be a better investment than an Institutional Repository?
  • Does anyone care now we have Facebook?

Wednesday, 15 December 2010

Myths about students -- and implications for Web design

Interesting post by Jakob Nielsen which claims that usability research undermines some prevailing myths about students.

Myth 1: Students Are Technology Wizards
Students are indeed comfortable with technology: it doesn't intimidate them the way it does some older users. But, except for computer science and other engineering students, it's dangerous to assume that students are technology experts.

College students avoid Web elements that they perceive as "unknown" for fear of wasting time. Students are busy and grant themselves little time on individual websites. They pass over areas that appear too difficult or cumbersome to use. If they don't perceive an immediate payoff for their efforts, they won't click on a link, fix an error, or read detailed instructions.

In particular, students don't like to learn new user interface styles. They prefer websites that employ well-known interaction patterns. If a site doesn't work in the expected manner, most students lose patience and leave rather than try to decode a difficult design.

Myth 2: Students Crave Multimedia and Fancy Design
Students often appreciate multimedia, and certainly visit sites like YouTube. But they don't want to be blasted with motion and audio at all times.

One website started to play music automatically, but our student user immediately turned it off. She said, "The website is very bad. It skips. It plays over itself. I don't want to hear that anymore."

Students often judge sites on how they look. But they usually prefer sites that look clean and simple rather than flashy and busy. One user said that websites should "stick to simplicity in design, but not be old-fashioned. Clear menus, not too many flashy or moving things because it can be quite confusing."

Students don't go for fancy visuals and they definitely gravitate toward one very plain user interface: the search engine. Students are strongly search dominant and turn to search at the smallest provocation in terms of difficult navigation.

Myth 3: Students Are Enraptured by Social Networking
Yes, virtually all students keep one or more tabs permanently opened to social networking services like Facebook.

But that doesn't mean they want everything to be social. Students associate Facebook and similar sites with private discussions, not with corporate marketing. When students want to learn about a company, university, government agency, or non-profit organization they turn to search engines to find that organization's official website. They don't look for the organization's Facebook page.

Friday, 3 December 2010

Show me the numbers

One of the offshoots of the Arcadia Project was the joint UL/CARET CULWidgets product, which wangled some JISC funding to "provide users with services appropriate to a networked world" in a widgetty/web services way.

Our two main production interfaces are the Cambridge Library Widget and the CamLib mobile web app, both soft-launched at the start of this term. This, the last day of term, seems a good time to look back and see how they've done.

Overall unique visitor numbers for the Widget are:



And, slightly more erratically, for CamLib:


A combined 4,489 unique visitors across the two. The services are mainly targeted at Undergraduates, of which Cambridge has c.12,000. Even assuming some crossover between the interfaces, and the likelihood that not all of them are undergrads, we're still looking at a significant proportion of our UG population (25-33%?)

Of course, these are just our unique visitors (i.e. distinct people who have visited the site) - total visits for the period are 12,284 and 2,559 respectively. Which shows that students are coming back to the interfaces again and again, not just taking a crafty peek. Monthly figures across both interfaces average around 3,000 unique users.

Our initial target was for 2,000 unique users in the first term, so we're running at well over double. Well done Widgets! And there are more to come.

Wednesday, 1 December 2010

Disruptive technologies in digitisation

Much of my fellowship has been taken up with examining three tech initiatives, all of which could be used in an on-demand process and could also be classed as disruptive. One is software, two are hardware.

Here is a bit more information ...

1) The Copyright calculator

Public Domain Calculators from Open Knowledge Foundation on Vimeo.



What is it?
A software development that assesses the copyright status of a creative work by looking at associated metadata. I've made an initial attempt to tie the Open Knowledge Foundation calculator into LibrarySearch, or new catalogue interface.

Why is it disruptive?
It can give the reader a useful indication of the copyright status of a book, allowing them to decide how they can re-use it. Its potentially useful as the first stage in a digitisation selection workflow, but also useful on its own. Its also an example of the commoditisation of a basic legal service.

What problems are there?
The be effective, the calculator needs author death-date information. Libraries only record this information when they wish to differentiate a name. Linked data tying a record into other sources of information could help overcome this


2) Kirtas book-scanner



What is it?
An automatic book-scanner. Turns pages using a vacuum equipped robot-arm and images pages with dual high-spec cameras

Why is it disruptive?
Books can be scanned and turned into PDF or other documents in a matter of hours, with minimal human interaction required.

What problems are there?
Its not cheap and still not 100% accurate. Its also a robot, so should not entirely be trusted.


3) Espresso book machine



What is it?
A photocopy sized book creation machine that does not require a printing-works to run it. Can print and bind a book in minutes.

Why is it disruptive?
Provides a library or a bookstore with a massive research collection/ back catalogue with none of the storage problems or overheads . Could have implications on acquisition, collection development and every part of library activity.

What problems are there?
As with the Kirtas, its not cheap, and limited in formats and outputs. And its a robot.

I'll have more to say on my project as I slog through write-up ...

Futurebook 2010

Whilst beginning to wrap up my fellowship (more in another post), I took time out to attend the FutureBook conference yesterday. Organised by the Bookseller, this conference brought together a number of industry leaders to highlight their successes and to raise awareness of issues they have faced in digital publishing. It was a fascinating day. For publishers and booksellers alike, it seems the digital revolution has finally arrived. Some highlights:

  • The Bookseller has conducted a wide survey of the sector to gauge opinion and attitude, with over 2,600 responses. This will be published soon.
  • One statistic was worth noting, when asked who will gain most from a rise in digital sales, respondents suggested readers, authors, publishers as the most, with booksellers and libraries rated last.
  • Publishers and booksellers had differing ideas regarding how quickly the change would occur. By the end of 2015, 2/3rds of publishers believed digital sales would account for anywhere between 8-50% of the market. Only just over half of the booksellers polled believed the same thing
  • Google will enter the online book retail market soon with Google Editions. Rather than tie themselves to a device, they are aiming for a platform agnostic browser and app based model, with all content remaining in the cloud rather than on-devices (although HTML 5 based local storage will be used) It will allow various online retailers and booksellers to build platforms around Google editions.
  • Tech-startups were suddenly seen as competition by publishers, at least in the app business. In response, much value was placed on publishers' knowledge of markets, talent and trends, as well as the curatorial process of commissioning and editing
  • Richard Mollet of the Publishers Association talked up the Digital Economy Bill. Formerly in the music industry, he noted that 'rights and copyright make the digital world go round', and argued that the bill was vital in explaining the damage illegal copying had on the creative sectors
  • Nick Harkaway, an industry commentator agreed in principle, but noted that enforcement so far had failed to deter illegal filesharing and DRM was no serious barrier to rights infringement. He urged publishers to keep people paying by offering serious innovation rather than simple digital recyling of print content
  • The academic book sector was well represented, with Wiley, CUP, OUP , Ingenta Connect and Blackwells Academic presenting. OUP gave an excellent talk on the changes required across an institution
  • We also had displays from Scholastic regarding the cross-media Horrible Histories series and looks at the editorial and creative processes behind booksellers' first steps into the world of mobile app development. Max Whitby from Touchpress showed off the Elements app, the next stage in the evolution of the coffee-table book
  • YouGov have a tablet track scheme looking at customer experiences of iPads and Kindle readers, which produced some interesting facts.
One over-arching trend that libraries can learn from relates to changes in the production and publishing processes. The phrase 'reflow-able text' was heard throughout the day, with publishers being urged to ditch PDF and print centric work flows in favour of granular xml-based marked up text that could be easily re-purposed for the next device of platform.

It seems to me that mainstream publishing is jumping over academic publishing on this one. Given the amount of on-line journal vendors that still insist on forcing PDF files down our throats, XML based delivery of more academic content could be of real use now to the consumer as well as the publisher.

One application of this approach was demonstrated, the Blackwells Academics custom textbooks service. This allows course-leaders to assemble all the material relating to a course into one bound volume which could then be sold on. Taking care of rights clearance, it also quite handily passed on the cost of printing course-packs to students!

Its still a great concept. Such a leap is only possibly by storing content in a normalised XML form, allowing it to be quickly pulled together to create new outputs.

Of course, we've been doing this in libraries with TEI and other transcription initiatives for some time now, but publishing at least is really taking the concept to heart, especially when faced with multiple devices and platforms to support. Post iPad, library digitisation projects will need to bear this delivery model in mind rather than relying upon image based delivery.

Tuesday, 23 November 2010

Analog tools

Workspace

Sometimes, you just can’t beat an olde-worlde paper notebook. Highly portable, great screen resolution, excellent, intuitive user interface and infinite battery life.

Only problem: it’s hard to back up. On the other hand, it’ll still be readable in 200 years. Which is more than can be said for any of my digital data.

Saturday, 20 November 2010

Digital Humanities

The New York Times has an interesting piece about the renewal of interest in the Digital Humanities.

The next big idea in language, history and the arts? Data.

Members of a new generation of digitally savvy humanists argue it is time to stop looking for inspiration in the next political or philosophical “ism” and start exploring how technology is changing our understanding of the liberal arts. This latest frontier is about method, they say, using powerful technologies and vast stores of digitized materials that previous humanities scholars did not have.


The article goes on to describe a few interesting projects. For example:


In Europe 10 nations have embarked on a large-scale project, beginning in March, that plans to digitize arts and humanities data. Last summer Google awarded $1 million to professors doing digital humanities research, and last year the National Endowment for the Humanities spent $2 million on digital projects.

One of the endowment’s grantees is Dan Edelstein, an associate professor of French and Italian at Stanford University who is charting the flow of ideas during the Enlightenment. The era’s great thinkers — Locke, Newton, Voltaire — exchanged tens of thousands of letters; Voltaire alone wrote more than 18,000.

“You could form an impressionistic sense of the shape and content of a correspondence, but no one could really know the whole picture,” said Mr. Edelstein, who, along with collaborators at Stanford and Oxford University in England, is using a geographic information system to trace the letters’ journeys.

He continued: “Where were these networks going? Did they actually have the breadth that people would often boast about, or were they functioning in a different way? We’re able to ask new questions.”

One surprising revelation of the Mapping the Republic of Letters project was the paucity of exchanges between Paris and London, Mr. Edelstein said. The common narrative is that the Enlightenment started in England and spread to the rest of Europe. “You would think if England was this fountainhead of freedom and religious tolerance,” he said, “there would have been greater continuing interest there than what our correspondence map shows us.”

Saturday, 13 November 2010

Hacking the Library -- ShelfLife@Harvard

What is Shelflife?
Shelflife is web application that uses what libraries know (about books, usage and comments) to allow researchers and scholars to access the riches of Harvard’s collections through a simple search.

Researchers will be able to access, read about, and comment on works using common social net- work features. ShelfLife will bring Harvard results to the forefront of the research process, allowing users to easily access and explore our vast collections.
What makes it unique?

Shelflife is designed to help you find the next book. Each search will retrieve a unique web page providing key information about the thing searched, including basic information, fluid links to related neighborhoods, and analytic data about use, all presented in a clean graphical format with intuitive navigation with discoverability in mind.


From the Harvard Library Innovation Lab. The site provides no information about ShelfLife beyond the above, but Ethan Zuckerman, who's a Berkman Fellow at the moment, has a useful blog post reporting a presentation by David Weinberger and Kim Dulin, who co-direct the project.
Libraries tend to be very knowledgeable about what they hold in their collections. But they’re much less good about helping people discover that information. There are few systems like Amazon or Netflix recommendations that help scholars and researchers discover the good stuff within libraries. Dulin argues that librarians have been pretty passive in the face of new technology – they’ve purchased fairly primitive systems and had to buy back their content from the companies who build those systems.

Researchers tend to start with Google, Dulin tells us. They might move to Google Books or Amazon to find out more about a specific book. And perhaps a library will come into play if the book can’t be downloaded or purchased inexpensively. Libraries would like to move to the front of that process, rather than sitting passively at the end. And lots of libraries are trying to take on this challenge – new librarians often come out of school with skills in web design and application development.

The Lab hopes to bring fellows into the process, much as Berkman does. It works to build software, often proof of concept software. And innovation happens on open systems and standards, so libraries and other partners can adopt the technology they’re developing.

Two major projects have occupied much of the Lab’s time – Library Cloud and ShelfLife, both of which Weinberger will demo today. There are smaller applications under development as well. Stackview allows the visualization of library stacks. Check Out the Checkouts lets us see what groups of users are borrowing – what are graduate divinity students reading, for instance. And a number of projects are exploring Twitter to share acquisitions, checkouts and returns.

Weinberger explains that ShelfLife is built atop Library Cloud, a server that handles the metadata of multiple libraries and other educational institutions and makes that metadata available via API requests and “data dumps”. Making this data available, Weinberger hopes, will inspire new applications, including ones we can’t even imagine. ShelfLife is one possible application that could live atop Library Cloud. Other applications could include recommendation systems, perhaps customized for different populations (experts, versus average users, for instance.)


Turns out the ShelfLife is in a pre-Alpha state of development. The metaphor behind it is the "neighbourhood" -- i.e. clusters that a given book might sit within.
We see a search for “a pattern language”, referring to Christopher Alexander’s influential book on architecture and urban design. We see a results page that includes a new factor – a score that indicates how appropriate a title is for the search. We can choose any result and we’ll be brought into “stack view”, where we can see virtual books on a shelf as they are actually sequenced on the physical shelf. Paul explains that it’s actually much more powerful than that – many books at Harvard are in a depository and never see the light of a shelf. And many colelctions have their own special indices – the virtual shelf allows a mix of the Library of Congress categories with other catalogs.

The system uses a metric called “shelfrank” to determine how the community has interacted with a specific book. The score is an aggregate of circulation information for undergraduates, graduates and faculty, information on whether the book has been assigned for a class, placed on reserve, put on recall, etc. That information exists in Library Cloud as a dump from Harvard’s HOLLIS catalog system – in the future, the system might operate using a weekly refresh of circulation data. The algorithm is pretty arbitrary at this point – it’s more a provocation for discussion than a settled algorithm.


Ethan reports some of the Q&A and generally does a great job of writing up the event. His post is worth reading in full.

A systems view of digital preservation

The longer I've been around, the more concerned I become about long-term data loss -- in the archival sense. What are the chances that the digital record of our current period will still be accessible in 300 years' time? The honest answer is that we don't know. And my guess is that it definitely won't be available unless we take pretty rigorous steps to ensure it. Otherwise it's posterity be damned.

It's a big mistake to think about this as a technical problem -- to regard it as a matter of bit-rot, digital media and formats. If anything, the technical aspects are the trivial aspects of the problem. The really hard questions are institutional: how can we ensure that there are organisations in place in 300 years that will be capable of taking responsibility for keeping the archive intact, safe and accessible?

Aaron Schwartz has written a really thoughtful blog post about this in which he addresses both the technical and institutional aspects. About the latter, he has this to say:

Recall that we have at least three sites in three political jurisdictions. Each site should be operated by an independent organization in that political jurisdiction. Each board should be governed by respected community members with an interest in preservation. Each board should have at least five seats and move quickly to fill any vacancies. An engineer would supervise the systems, an executive director would supervise the engineer, the board would supervise the executive director, and the public would supervise the board.

There are some basic fixed costs for operating such a system. One should calculate the high-end estimate for such costs along with high-end estimates of their growth rate and low-end estimates of the riskless interest rate and set up an endowment in that amount. The endowment would be distributed evenly to each board who would invest it in riskless securities (probably in banks whose deposits are ensured by their political systems).

Whenever someone wants to add something to the collection, you use the same procedure to figure out what to charge them, calculating the high-end cost of maintaining that much more data, and add that fee to the endowments (split evenly as before).

What would the rough cost of such a system be? Perhaps the board and other basic administrative functions would cost $100,000 a year, and the same for an executive director and an engineer. That would be $300,000 a year. Assuming a riskless real interest rate of 1%, a perpetuity for that amount would cost $30 million. Thus the cost for three such institutions would be around $100 million. Expensive, but not unmanageable. (For comparison, the Internet Archive has an annual budget of $10-15M, so this whole project could be funded until the end of time for about what 6-10 years of the Archive costs.)

Storage costs are trickier because the cost of storage and so on falls so rapidly, but a very conservative estimate would be around $2000 a gigabyte. Again, expensive but not unmanageable. For the price of a laptop, you could have a gigabyte of data preserved for perpetuity.

These are both very high-end estimates. I imagine that were someone to try operating such a system it would quickly become apparent that it could be done for much less. Indeed, I suspect a Mad Archivist could set up such a system using only hobbyist levels of money. You can recruit board members in your free time, setting up the paperwork would be a little annoying but not too expensive, and to get started you’d just need three servers. (I’ll volunteer to write the Python code.) You could then build up the endowment through the interest money left over after your lower-than-expected annual costs. (If annual interest payments ever got truly excessive, the money could go to reducing the accession costs for new material.)

Any Mad Archivists around?


Worth reading in full.

LATER: Dan Gillmor has been attending a symposium at the Library of Congress about preserving user-generated content, and has written a thoughtful piece on Salon.com about it.

The reason for libraries and archives like the Library of Congress is simple: We need a record of who we are and what we've said in the public sphere. We build on what we've learned; without understanding the past we can't help but screw up our future.

It was easier for these archiving institutions when media consisted of a relatively small number of publications and, more recently, broadcasts. They've always had to make choices, but the volume of digital material is now so enormous, and expanding at a staggering rate, that it won't be feasible, if it ever really was, for institutions like this to find, much less, collect all the relevant data.

Meanwhile, those of us creating our own media are wondering what will happen to it. We already know we can't fully rely on technology companies to preserve our data when we create it on their sites. Just keeping backups of what we create can be difficult enough. Ensuring that it'll remain in the public sphere -- assuming we want it to remain there -- is practically impossible.


Dan links to another thoughtful piece, this time by Dave Winer. Like Aaron Schwartz, Dave is concerned not just with the technological aspects of the problem, but also with the institutional side. Here are his bullet-points:

1. I want my content to be just like most of the rest of the content on the net. That way any tools create to preserve other people's stuff will apply to mine.

2. We need long-lived organizations to take part in a system we create to allow people to future-safe their content. Examples include major universities, the US government, insurance companies. The last place we should turn is the tech industry, where entities are decidedly not long-lived. This is probably not a domain for entrepreneurship.

3. If you can afford to pay to future-safe your content, you should. An endowment is the result, which generates annuities, that keeps the archive running.

4. Rather than converting content, it would be better if it was initially created in future-safe form. That way the professor's archive would already be preserved, from the moment he or she presses Save.

5. The format must be factored for simplicity. Our descendents are going to have to understand it. Let's not embarass ourselves, or cause them to give up.

6. The format should probably be static HTML.

7. ??

Sunday, 7 November 2010

Put not your faith in cloud services: they may go away

From John Dvorak:
I have complained about the fly-by-night nature of these companies for years, but my concern now seems misplaced. I was concerned about operations that you depend on for deep cloud services. This means complex programs running on the cloud with no real alternative. Over time, I've tended to see these companies as more stable than the "Use our free service. You won't regret it!" model.

I was taken to task by numerous vendors who kept telling me that I was full of crap, because cloud services are professionally managed, and nobody could do the job—whatever the job was—better than a room of pros. With the cloud, the pros would also keep the data safe.

Yeah, until they were all laid off, and the service shut down!

Now here's the problem I am experiencing second-hand. The audio podcast I do with Adam Curry, the No Agenda Show (Google it), has been using Drop.io to store podcast album cover images for convenience. They will all be destroyed, as well as the accumulation of links, tips, curiosities, and other valuable information, in the next few weeks.

Looking back on the idea of using this service, I didn't fully consider the ramifications of its discontinuance despite my skepticism about cloud services in general. You know, this was just a lot of weird stuff thrown into a bin. But once it was discontinued, it was apparent what you are left with: dead links.