Swadesh List data now re-enabled in Rosetta Internet Archive Collection

Swadesh list for the Puoc language in the International Phonetic Alphabet

In the 01950s, American linguist Morris Swadesh, as part of his overarching vision of a quantitative method for determining language relationships on a global and multimillenial scale, developed a set of one hundred words found to be unusually stable across time and language boundaries. Swadesh hypothesized that words like “fire,” “moon,” “mother” and “bone,” common to human experience, were far less likely to change or be substituted with words borrowed from other dialects or languages. The 100 word “Swadesh list” (sometimes up to 207, depending on the variety of the list used) is now widely collected in linguistic field research, and functions as a kind of universal linguistic fossil. With careful study, these lists can reveal ancient language relationships and processes of linguistic change typically obscured by centuries-long processes of evolution and borrowing. As familiar examples, such processes transformed Chaucer’s English into modern English and Latin into the modern Romance Languages.

In 02004, The Rosetta Project undertook a National Science Foundation funded project to increase both the size and utility of its long-term multilingual archive and at this time added a large number of Swadesh lists to its collection. Lexical database archivists Tim Usher and Paul Whitehouse contributed original research (Tim Usher’s 02002 Indo-Pacific database and Paul Whitehouse’s 02002 Australian and New Guinea database were central among the additions) and also brought in outside resources, including Darrell Tryon’s Comparative Austronesian Dictionary (01995), George Starostin’s Dravidian database, and Ilya Peiros’ Mon Khmer database. In many of these cases, as with the Usher and Whitehouse collection, the 100-200 term Swadesh lists were a subset of a larger lexical data collection project. Despite the Swadesh list’s limitation in size compared with a resource like a dictionary, a large collection of the same material in many different languages is useful as a parallel dataset for cross-linguistic comparison.

This collection of Swadesh lists was included as a parallel data set among the documents micro-etched on the Rosetta Disk, a physical copy of The Rosetta Project’s long-term linguistic archive created in 02008. And for a period of time, the lists were available on The Rosetta Project’s website via an interactive tool which allowed visitors to view and compare lexical items in over a thousand languages and also contribute their own lexical data. But as the Rosetta Project site evolved and the structure of serving environments changed, this tool became technologically obsolete. While there was (and remains) no lack of storage space for the lists, there was a critical lack of what Long Now board member Kevin Kelly calls “movage.”

“Movage,” says Kelly, means transferring the material to current platforms on a regular basis — that is, before the old platform completely dies, and it becomes hard to do. This movic rhythm of refreshing content should be as smooth as a respiratory cycle — in, out, in, out. Copy, move, copy, move.” And it is movage, not storage, says Kelly, that is critical to keeping information alive: “The only way to archive digital information is to keep it moving.” In other words, simply storing data isn’t enough to ensure its longevity; it must be copied, moved, and made redundant. And not just once or twice — indefinitely. Kurt Bollacker, Long Now Foundation Digital Research Director, adds: “[b]ecause any single piece of digital media tends to have a relatively short lifetime, we will have to make copies far more often than has been historically required of analog media. Like species in nature, a copy of data that is more easily “reproduced” before it dies makes the data more likely to survive.” [1]

Since the 02004 iteration of the Swadesh list program, The Rosetta Project has launched a comprehensive migration of all of its data to The Internet Archive, a free online digital library founded in 01996 with over 4 petabytes of storage. The Internet Archive exemplifies the paradigm shift in the field of information preservation from storage to movage: users of the site can upload any document they have permission to distribute to the site for free, where anyone with access to the internet can then download it to their own machine. Thousands of downloads are made every day from Internet Archive servers by users all over the world: early “movage” on a massive scale.

After a long process of unraveling and decoding the Swadesh list data, which had fallen victim to rapid changes in character encoding and database standards, The Rosetta Project has now moved the collection of 1,235 Swadesh lists into The Internet Archive. Recognizing the substantial merit and long-term advantages of the movage model and its successful early implementation by The Internet Archive, our goal is for the lists to have a long, useful, and redundant residence there.

The relocation of the Swadesh lists is also the first step of The Rosetta Project’s latest undertaking, The 300 Languages Project. Source materials collected for The 300 Languages Project, whose aim is to address a need for highly-structured linguistic resources in the world’s 300 most widely-spoken languages, will be stored at The Internet Archive with the rest of The Rosetta Project collection.

Was the 5-to-6-year period the Swadesh list data spent in the darkness unusual? According to Kelly, not at all: “We don’t know what the natural movage respiration cycle is for digital media yet since it is still very new,” says Kelly, “but I suspect the cycle is much shorter than we think. I would guess it is 5 years. No matter what digital format you have your precious [data] stored on, you should expect to move it onto new media in five years — and five years after that forever!”

Ideas

Swadesh List data now re-enabled in Rosetta Internet Archive Collection