We sat down with David Willis, currently the Jesus Professor of Celtic at the Faculty of Linguistics, Philology and Phonetics, University of Oxford to talk about the Parsed Historical Corpus of the Welsh Language (PARSHCWL), a project that started with a gap in the digital humanities landscape and has since grown into an international research collaboration spanning Cambridge, Oxford, and Berlin.
For those not familiar with the project, how would you describe what you were trying to build and why it mattered?
The point of the project was to build a corpus of historical Welsh texts that was annotated in ways that we and other researchers could use. Historical linguistics has always been based on looking at historical texts. In the past, you used to just read them. These days, you search them with a computer, and we have increasingly sophisticated ways of doing that. For example, you want to be able to search for grammatical structures not just particular words. So, the point of having a parsed corpus is not just to digitise the characters of a text, but to enrich it with information about grammatical structure that you can then use in your searches. And this information may be used to answer other research questions you might have later on.
And this is the really exciting part. Of course you have to build the corpus first, but what we really wanted to do was be able to ask the questions that researchers in historical linguistics care about, like how language changes.
What was missing before PARSHCWL, and why did it matter?
We simply did not have this resource for Welsh, and it was becoming pretty standard for other languages. Parsed historical corpora—treebanks, as they are called—already existed for English, German, French, Icelandic, Portuguese and others. Welsh had none. And that absence was a real obstacle to research on the language.
You have described Welsh as typologically unusual. What does that mean for a non-specialist, and why does it make a historical corpus of Welsh particularly valuable?
What we meant by typologically unusual is that Welsh has a few grammatical features that are not common in European languages. It is a verb-initial language, that is, the standard word order in a main clause is verb, subject, object. Only a handful of European languages do that, and under 10% of the world's languages. Welsh also has initial consonant mutation, where the first sound of a word changes according to its grammatical environment. That can cause difficulties, as a lot of automated procedures work by looking at the first few letters of a word, so if the first letter can change, that makes things considerably harder.
Because Welsh has these typologically unusual properties, it is important to include it in work on language change. There is a danger that historical linguistics focuses on languages that are very similar to each other. The main European languages may seem very different from the inside, but compared to the world's languages they are really quite similar. A broader typological spread is good for the field.
The project brought together three departments: Linguistics, German, and ASNaC. How did those different perspectives shape the work?
The German perspective was useful because these corpora have already been developed for a lot of Germanic languages. So that input helped us compare what we were doing to what had been done elsewhere and brought familiarity through those existing corpora. ASNaC brought more philological knowledge of the texts themselves, that is, what you might find in a medieval Welsh manuscript, what problems you might encounter. And Linguistics brought an awareness of needing rich historical data, and of what you are ultimately going to use that data for: to answer theoretically interesting questions.
What does it actually mean to parse a historical text? What are you adding to the raw words on the page?
We are trying to do two things. The first stage is part-of-speech tagging: for every word, you want to say whether it is a noun, an adjective, a pronoun, and with a bit more detail than that: is it a plural noun, a singular noun, a first-person pronoun, and so on. Then, once you have done that, you try and create syntactic structure. We were using constituency parsing, so we try to draw a tree that represents the grammatical structure of each sentence.
Historical Welsh orthography is described as highly inconsistent. What challenges does that create?
A lot of these automated procedures have been developed for modern English, as they are trained on modern English newspapers, essentially. The problem with a medieval language is that there is no standardised orthography. Even a simple word could have thirty different spellings. And we do not have enormous amounts of text, so you cannot simply rely on pattern recognition. The initial consonant mutation makes it worse, because it is not just that the orthography is variable; the initial consonant can also change for perfectly good grammatical reasons, and that adds even more variation to contend with. So there’s a lot more difficulty working out the correct part of speech. It's mainly for part of speech tagging that orthography is a problem. Once you've got that, you're okay.
The CLS Incubator Fund provided two grants totalling £5,249 — a first round of £2,500 in November 2016, and a second of £2,749 in March 2017. The money went almost entirely on employing Marieke Meelen, then a recently completed PhD researcher who had developed the initial part-of-speech tagging work on the Mabinogi texts, to extend and deepen that work on new material. But perhaps more than the text processing itself, what the funding bought was time to do something “unglamorous” and essential: write the manual of conventions.
The articles coming out of the project describe the conventions developed for the corpus. How central was the seed funding to producing those?
Absolutely central. In these articles we essentially wrote up what we did during the incubator fund projects. They explain what conventions we developed for the corpus and were then used to apply for follow-on funding. These conventions are really important. We could not just copy someone else's conventions, because we were working with a language that had not had this done before. And so establishing those conventions was the most difficult part, one we spent a lot of time doing. The first project was really to start developing the manual of conventions. We could not even begin processing text until we had done that, because the first sentence you look at has to fully implement the manual.
What happened after the CLS funding ended? Where did the project go from there?
So after the incubator funding, we put together an application to the British Academy/Leverhulme Trust Small Research Grant scheme, which is a little bigger than the incubator fund. By that point we knew enough: the incubator fund had let us put together the broad outline of the workflow and the basic parameters of the conventions. The British Academy grant allowed us to fund a workshop, bring more people in, and get to a point where the conventions were solid enough to build on.
And then we wanted to actually apply the corpus to real research questions in historical linguistics, not just keep building infrastructure. So, we looked for collaborators. We had been interested in the question of how Welsh lost null subjects. Specifically, Welsh used to allow sentences without an explicit subject, and it no longer does, and we wanted to trace when and how that change happened. We found partners in Berlin: Roland Meyer at Humboldt University, a Slavist I knew from previous work, who had done related work on the history of null subjects in Polish and Russian. Together we put together a much broader project, the history of pronominal subjects in the languages of Northern Europe, which was funded by the AHRC and the German Research Foundation (AHRC-DFG) jointly. That employed a researcher in Berlin working on Slavic languages, and a researcher working on Celtic, using PARSHCWL as the Welsh component of a larger comparative project.
Looking back, what did the CLS funding make possible?
We wouldn’t be where we are. We would not have been able to put in the final application, because we would not have known enough to put it together. The incubator fund let us establish the groundwork: the workflow, the parameters, the conventions. Without that, everything that came after would not have been possible.
What would you tell a researcher considering applying to the Incubator Fund?
I would not hesitate to apply. The great thing about the incubator fund is that it is not very bureaucratic, you only have to fill out quite a simple form. It is much easier than writing a full AHRC or Leverhulme application. That matters, because the whole point of an incubator fund is to give an idea the room to grow. You do not need to have everything worked out in advance, that is precisely what the funding is for. A few thousand pounds can make a lot of difference to your ability to put other things together. So my advice would simply be: if you have an idea that could benefit from a small interdisciplinary push, do not overthink it. Put in the application.
PARSHCWL continues to grow. The corpus is now being used to answer the questions it was always designed to enable, tracing how Welsh grammar changed over centuries of texts. The conventions developed with CLS seed funding underpin two peer-reviewed journal articles in the Journal of Celtic Linguistics and the Journal of Historical Syntax. What began with a simple application and a few thousand pounds of seed funding has, in the years since, helped establish Welsh as a full participant in the international landscape of historical corpus linguistics, and set in motion a chain of collaboration that now spans two countries and several research councils.