Archiving – Page 3 – Endangered Languages and Cultures

Could DNA be the future of digital preservation?

24 January 2013 by Aidan Wilson

Genetic scientists in Britain overnight, have successfully demonstrated the data-storage potential of DNA, as explained in The Conversation today. In a proof-of-concept experiment, a string of DNA with a physical size around that of a grain of dust, was encoded with an MP3 file of the ‘I have a dream’ speech of Dr Martin Luther King, … Read more

Counting Collections

30 April 201329 November 2012 by Nick Thieberger

As will be clear to regular readers of this blog, we are concerned here to encourage the creation of the best possible records of small languages. Since much of this work is done by researchers (linguists, musicologists, anthropologists etc.) within academia, there needs to be a system for recognising collections of such records in themselves as academic output. This question is being discussed more widely in academia and in high-level policy documents as can be seen by the list of references given below.

The increasing importance of language documentation as a paradigm in linguistic research means that many linguists now spend substantial amounts of time preparing corpora of language data for archiving. Scholars would of course like to see appropriate recognition of such effort in various institutional contexts. Preliminary discussions between the Australian Linguistic Society (ALS) and the Australian Research Council (ARC) in 2011 made it clear that, although the ARC accepted that curated corpora could legitimately be seen as research output, it would be the responsibility of the ALS (or the scholarly community more generally) to establish conventions to accord scholarly credibility to such products. Here, we report on some of the activities of the authors in exploring this issue on behalf of the ALS and discuss issues in two areas: (a) what sort of process is appropriate in according some form of validation to corpora as research products, and (b) what are the appropriate criteria against which such validation should be judged?

“Scholars who use these collections are generally appreciative of the effort required to create these online resources and reluctant to criticize, but one senses that these resources will not achieve wider acceptance until they are more rigorously and systematically reviewed.” (Willett, 2004)

Announcing Paradisec’s new catalogue

24 October 2012 by Aidan Wilson

Over the last year or so, the Paradisec team, in collaboration with software developers Robot Parade, Silvia Pfeiffer and John Ferlito, have been working on the development of a replacement to our ageing catalogue and database systems and a couple of weeks ago, this work culminated in the release of the new catalogue. There are … Read more

PARADISEC’s ‘Data Seal of Approval’

19 September 201219 September 2012 by Nick Thieberger

As we approach our tenth year of operation, it is gratifying that PARADISEC has achieved this seal of approval (DSA), based on 16 criteria (listed below, and see how we meet these criteria here: https://assessment.datasealofapproval.org/assessment_75/seal/html/). We have been a five-star Open Language Archives Community repository for some time, which also means that we are one of the 1800 archives whose catalog and metadata conform to the Open Archives Initiative standards, but the DSA looks more broadly at the whole process of the repository, from accession of records, through their description and curation and to disaster management. This is important for our depositors to know as they can be sure that their research output is properly described and curated, and can be found using various search tools, including google, but more specifically the Australian National Data Service, OLAC and the WorldCat, and also the aggregated information served in the Virtual Language Observatory.

Bursting through Dawes (2)

31 August 201231 August 2012 by David Nash

Further to my last post, I’ve read on, and my disappointment has only deepened at the treatment of the Sydney Language in Ross Gibson’s 26 views of the starburst world.

Think about the notes you made when you were getting into learning an undocumented language … Imagine they get archived and in a century or two someone looks through them and tries to work out what was going on when you made the notes. With only shreds of metadata and general knowledge of the historical period to go on, the future reader makes inferences from the content. Could a cluster of words in one of your vocabulary lists point to a hunch you were checking? Or a sequence of illustrative sentences could be the skeletal narrative of a memorable experience shared with your teachers.

Charting Vanishing Voices: A Collaborative Workshop to Map Endangered Oral Cultures

4 July 2012 by Nick Thieberger

A two-day conference titled ‘Charting Vanishing Voices: A Collaborative Workshop to Map Endangered Oral Cultures’ ran on June 29/30 in Cambridge, UK. Organised by the World Oral Literature Project, the conference brought together a range of ‘scholars, digital archivists and international organisations to share experiences of mapping ethno-linguistic diversity using interactive digital technologies.’ A discussion … Read more

Technology and language documentation: LIP discussion

27 June 2012 by Lauren Gawne

Lauren Gawne recaps last night’s Linguistics in the Pub, a monthly informal gathering of linguists in Melbourne to discuss topical areas in our field.

This week at Linguistics in the Pub it was all about technology, and how it impacts on our practices. The announcement for the session briefly outlined some of the ways technology has shaped expectations for language documentation:

The continual developments in technology that we currently enjoy are inextricably connected to the development of our field. Most would agree that technology has changed language documentation for the better. But while nobody is advocating a return to paper and pen, most would concur that technology has changed the way we work in unexpected ways. The focus is usually on the materials we produce such as video, audio and annotation files as well as particular types of computer-aided analysis. In a recent ELAC post, ‘Hammers and nails‘ Peter Austin claims that metadata is not what it was, in the days of good old reel-to-reel tape recorders. The volume of comments suggests that this topic is ripe for discussion. This session of Linguistics in the Pub will give us a chance to reflect on how our practices change with advances in technology.

There are a (very) few linguists who advocate that researchers should go to the field with nothing beyond a spiral-bound notebook and a pen, though no one at the table was quite willing to go that far; all of us, it seems, go to the field with a good quality audio recorder at the very least. Without the additional recordings (be they audio or video) the only output of the research becomes the final papers written by the linguist, which are in no way verifiable. The recording of verifiable data, and the slowly increasing practice of including audio recordings in the final research output are allowing us to further stake our claim as an empirical and verifiable field of scientific inquiry. Many of us shared stories of how listening back to a recording that we had made enriched the written records that we have, or allow us to focus on something that wasn’t the target of our inquiry at the time of the initial recording. The task of trying to do the level of analysis that is now expected for even the lowliest sketch grammar is almost impossible without the aid of recordings, let alone trying to capture the subtleties present in naturalistic narrative or conversation.

ELAR cracks a ton

25 June 2012 by admin

The Endangered Languages Archive (ELAR) at SOAS reaches an important milestone this week when our 100th deposit goes online. We will be working on a further 10 deposits and doing additional curation work on those currently online over the next two months. ELAR now has 4 terabytes (4,000 gigabytes — double that I reported in … Read more

‘e’s a diamond (jubilee) geezer, innit

6 June 2012 by admin

Most of the UK seems to have been distracted over the past few weeks (and especially over the four-day long weekend that is just now drawing to an end) by the celebrations surrounding the Diamond Jubilee of Queen Elizabeth II. Not so the hard working team at the Endangered Languages Archive (ELAR) at SOAS who … Read more

Australian Aboriginal Language Materials in ELAR

19 May 2012 by admin

If you are interested in Australian Aboriginal languages you might like to take at look at the growing number of collections of audio, video and text materials that are now available in the ELAR archive. Currently there are six online collections (comprising almost 900 file bundles) for languages from northern Australia, with one more from … Read more