Contributors

Showing posts with label Google. Show all posts
Showing posts with label Google. Show all posts

Saturday, June 18, 2011

Hard Drive, Catastrophe-Proof Bunker or Peat Bog?

   A couple of interesting items have appeared in the news this week relating to the interests of the people who frequent this blog. The first is that the Internet Archive is setting up a facility to preserve paper books that they have digitised. This relates to some issues covered in the posting Reprise on Google, e-Books, Copyright and All That Jazz. The idea is to have one copy of everything they can get their hands on, not for regular consultation, but as a "seed bank" if needed to check or resurrect digital copies. Take a look at the original article by Brewster Kahle Why Preserve Books? The New Physical Archive of the Internet Archive. There has already been considerable commentary on this posting and some on other blogs, some of it thoughtful, some of it a little nutty. I was slightly intrigued by the putative author who thought that archiving his book would be a breach of copyright. It's going to get a bit rough when authors come knocking on your door to see what you have done with their books.
   It was also of interest to note that the Internet Archive preserves old hard drives which contain digitised material, as well as microfilms. Somewhere along the line, I guess we will find out which of these media survives the best.
   The process was evidently inspired by hearing that some libraries were culling their collections after books had been digitised by Google. Now libraries have always culled, but I guess there might be a critical mass developing out there. It is about one good generation since a huge expansion in universities, and their libraries, and there is an ongoing expansion in the number of books, as well as journals, published. We are rarely presented with the definitive work on any subject these days. Decisions must be made about what to preserve and how to preserve it, and we all have our favourite causes. Trashing perfectly good old books always seems like murder. I mean, maybe everybody threw out their copy of A User's Guide to CP/M, like I did not so long ago.
   Meanwhile, in Dublin, the National Museum has just put on display an 8th century psalter found in a peat bog in 2006 and finally conserved so that it can be displayed. The museum has an article on the find and its conservation The Faddan More Psalter. Click on the PDF file link for more information on the conservation and some photographs. The Independent has an article about the item going on display Public gets first look at ancient book of psalms. Now while it seems that a peat bog is not the ideal conservation medium for ancient books, it is better than a damp cellar, inflammable library or bug ridden attic. The book is severely damaged, but it is still there and evidence for ancient Irish written culture. But why was it in a peat bog at all? Is it possible that it may have been culled from a monastic library centuries ago???
   There is a story that Gerald of Wales, the 12th century cleric and traveller through Britain's Celtic realms, saw the Book of Kells. He certainly saw and described an ancient book which he admired, but nothing about it suggests that it was the Books of Kells. Rare finds like this latest remind us that there were many wondrous things that have been lost.

Sunday, May 29, 2011

Reprise on Google, eBooks, Copyright and All That Jazz

  An enigmatic personage by the name of Dr Beachcomber has sent me an email with a link to his posting Google Burns the Library at Alexandria. He has included my reply as a comment on his blog, so I am returning the compliment by referencing him here. Is this what you call some kind of hippy blog-in?
  While I have mildly chastised him for over dramatics in headline writing, the books not actually being burned as a result of having been digitised, there is a issue of concern regarding the quality of scanned digital editions, and another issue brought up by another commenter on the recopyrighting of material already in the public domain as a result of it being reprinted or republished. There is also the very tricky issue of the destruction of original printed or written material after it is digitised.
  Taking the last first (Hey, I'm in Australia, we are upside down here!), I was many years ago doing a research project which involved examining museum records and objects. Now museum curators have a habit of updating their records when they think that a person looking at them is some kind of expert and they ask them questions about things. For historical reasons, I wanted to know what the original records said about the objects. With the old handwritten cards and registers, it was possible to separate the original records from later annotations, and even to work out who had made the annotations and when. Only one museum had an electronic catalogue at that time (1991), and they were quite disappointed that I actually wanted to look at their tatty old paper records. I guess the question is, how many old paper backups do we need to keep for safety? The same applies to books.
  On the second issue, I was told many years ago by a copyright legal bod in my university that it was legal for me to scan out of copyright visual material and republish it digitally, but it was illegal to reproduce digital scans from modern facsimile editions of out of copyright material. My only question about that is, how could anybody tell? At the moment the business interests are noisily defending ever increasing copyright restrictions, but the ready availability of copying and reproduction technology is going to make soup of that, and real soon. I suggest that if you have some favourite old, genuinely out of copyright, books in your particular area of interest or expertise, digitally reproduce them yourself, circulate them among your friends and colleagues, and loudly announce them as public domain.
  The quality of some of the old material scanned and placed in the public domain is an issue. Dr Beachcomber is determined that Internet Archive editions are better quality than those from Google, but I bet he has never spent three days printing a long book page by page from two separate Internet Archive scans, hoping that the pages missing from the two editions do not actually coincide at any point. The end result was a largely black and white edition with occasional colour pages, none of which had bookmarkable or cut and pastable text as they were simply image scans of pages. And the Kindle editions are similarly unnavigable and messily formatted. And the text only versions are unformatted to illegibility and full of OCR errors. But apart from that they're alright. I suspect that there is just some degree of luck with the digitisation of particular works, and how carefully they have been done.
  I have touched on these issues in earlier posts, Eeee! Books, and Scribes, Copyright, Crime and Google, with a short note at the end of Horrible Old Handwriting. I guess the whole issue is just not going to go away real soon.
   The whole issue of preservations of books and text is, of course, not new, but there are so many texts to preserve these days. We have almost no original Roman era texts of the Latin Classics, because they were written on papyrus rolls which fell to bits. These works are mainly preserved from much later copies in vellum codices, much more durable, produced by Christian monks. The thought of these celibate ascetics solemnly copying down the erotic poetry of Ovid and the like is always good for a giggle, but they did. There have even been conspiracy theories that the monks actually forged all the Latin Classics. I doubt it, but how much did they edit, correct, annotate and standardise these texts? Perhaps Cicero or Livy might be surprised to discover what we think they had written.
Postscript: With apologies to Dr Beachcomber, after rechecking, it seems that the download I had such trouble with was a Google scan, although I accessed it through the Internet Archive. It was one of a large set uploaded by one tpb, who seems to be a very messy worker. Perhaps I was dead unlucky, because there appears to be another edition of the same book available through the Internet Archive which is not from Google, so at least they are not claiming a monopoly for their grotty scans.

Thursday, June 03, 2010

Scribes, Copyright, Crime and Google

Now I know I do like to rabbit on sometimes about the continuities and discontinuities in written communication in the middle ages and today, but during the course of the last day or so I have truly found myself, like Alice, down the rabbit hole and behind the looking glass. It all started when I found a link to an interesting old French paleography book published in 1892.
In the days of medieval scribes, there was no such thing as copyright. No sooner had an author put away his quill than everybody was free to transcribe his words, paraphrase them and incorporate them into new contexts. It was only the industrial production of books which set up the conditions for protection of authors and publishers, and then not for some centuries. The purpose of copyright was not to inhibit the dissemination of words, but to encourage them by protecting the investment of those who had set up print runs of books. Authors got paid royalties, so they didn't have to rely on the patronage of kings and magnates in order to eat.
Many interesting books published long ago are no longer available, often because they are only interesting to a small number of people, but interesting nonetheless. Google has been collaborating with some very eminent libraries to make these available again through Google books, but there is a catch. Copyright laws are not the same the whole world over. Rather than try to untangle the mess, Google has simply made certain books unavailable in full text to countries outside America if they have been published between around 1870 to the 1920s, whatever their actual copyright status. This was the case with this old paleography book, written in France, which I was trying to access from Australia.
Trolling around the web to resolve this issue, I discover that one suggestion is to obtain a free proxy in America, so that Google thinks that is where you are. This is very easy. There are hundreds of them, and they make up new ones every day as the old ones are shut down. Furthermore, they advertise this service in terms of not allowing your web surfing to be tracked, and enabling you to access sites banned in your school, workplace or country of residence. The sites have a tendency to have "up yours" or "in your face" kinds of names. I have also discovered that this is the very easy fix to our very stupid country's very stupid proposed mandatory internet filtering. So here I am, consorting with pornographers, gunrunners, terrorists, clandestine Facebook users in the workplace and God knows who else, just in order to read a very old academic book which actually is thoroughly out of copyright, here and everywhere else. Furthermore, very respectable archivists and academics have shown me how to do it, because it is not illegal to read or download this book. It's just the kind of company you have to keep in order to do it.
Of course, there was a catch. The proxy would not allow me to download the book as a pdf file because it was too big. Another proxy with a large download allowance was not actually in the USA. So the second suggestion was to see whether it was on the Internet Archive as a text. It was, but ..... if you clicked on the link to download the pdf, you got sent to .... Google books. I am now part way through the process of printing a large book one page at a time from the Internet Archive, because it is one of those reference type books that are useful to have to hand. It's got huge bibliographies and a large dictionary of Latin abbreviations. Yawn! But it is still cheaper than flying over to Harvard to look at it in the library from which it was Googled.
Isn't it about time that the publishing industry got over its paranoia, and there was some means of releasing elderly books of specialised interest without hysteria about copyright? If the book is out of print, it should be available. The big publishers are not going to have their sales of the next J.K. Rowling or Dan Brown knocked off by a few harmless oddballs downloading dictionaries of medieval Latin abbreviations. I'm sure that long deceased authors of specialist academic material would be fascinated if they could know that somebody still did want to read what they had written, just as I'm sure that even longer deceased medieval scribes would be fascinated and bemused by people rescuing scraps of their work from the bindings of later books and poring over them.
Meanwhile, if you don't hear from me again, you will know what has happened. "Knock, knock!! "Ello Sunshine, you're nicked! You've been using HideMyAss in order to download a little cursiva bastarda. Just come along with me, Madam."

Wednesday, August 15, 2007

Google and Link Lists

One of the oddities I find whenever I check the web stats for Medieval Writing is that the site receives a regular trickle of hits from link sites which are seriously out of date. The users might have found Medieval Writing, but presumably they have also encountered a large number of 404s in their travels. Trolling around Google recently to see if anything new had popped up in my field of interest, it became apparent that the link sites which came up near the top of the Google list were mostly very out of date, with many dud links. Presmably, having established their place near the top of the list, they stay there if people keep trying to use them. Google, despite what it says in its own publicity, does not measure significance or relevance, only clicks.
I guess the responsibility for keeping the web up to date lies with the users. If you haven't updated your link site since around 2000 and don't intend to real soon, please take it down.