Some questions that motivate my research:

  • How does language vary and change within and across communities?
  • How do people use language variation to signify community identity, affiliation, and difference?
  • How do we build and maintain tools that support answering these questions?
  • How do community practices and technologies shape participation and retention online?

Supporting the linguistics community through software

Modern sociolinguistic research relies on software to automate data extraction and analysis—and this software does not write itself. Researchers use tools that time-align transcripts with speech sounds, a process known as forced alignment, and extract acoustic measurements from the identified sounds. One widely used tool is Forced Alignment and Vowel Extraction (FAVE), a package originally developed in part to study Philadelphia English. Although FAVE supported important research, its aging codebase had become difficult to install, maintain, and incorporate into contemporary research workflows.

Working with Josef Fruehwald (University of Kentucky), I helped improve and stabilize FAVE as a developer and maintainer. Through tutorials and conversations with researchers about analysis methods and workflows, I identified a recurring usability problem: researchers increasingly wanted to incorporate FAVE into larger data pipelines, but the software had been designed primarily as a standalone application.

To support these workflows, I migrated the codebase from Python 2 to Python 3, modularized it for use within other programs, packaged it for distribution through PyPI, and implemented automated testing, documentation, and continuous integration. As researchers adopted the updated version, I responded to questions and issues through email and GitHub. This work made FAVE easier to install, reuse, and maintain while supporting more transparent and reproducible research workflows.

Participation and community health on Wikipedia

For more than a decade, I have contributed to Wikimedia projects, where recruiting and retaining editors remains a persistent challenge. During that time, I have seen how early interactions, community norms, and technical features shape whether readers become contributors. Drawing on my editorial experience and research background, I have studied how these factors affect newcomers’ ability to participate successfully.

Wikipedia is the encyclopedia that anyone can edit, but on some pages the barrier to entry is high. Administrators can lock articles against contributions from newer and unregistered editors. To understand what happens when these barriers are lowered, I reviewed 55 articles that had remained locked since before 2010 and selected seven to unlock. Over 31 days, I manually coded contributions according to editor experience, apparent intent, and constructive value. Some pages needed to be locked again, but others received constructive contributions from new editors. Even imperfect edits sometimes drew attention from experienced editors and prompted further improvements—a phenomenon described by early wiki communities as PageChurn. The findings showed that broad restrictions can prevent good-faith participation, while automated moderation tools can manage disruption with less exclusion. I presented this research at Wikimania 2021.

Many prospective editors assume that contributing to Wikipedia means writing an entire article. During editor trainings and outreach events for linguists and other subject-matter experts, I repeatedly saw this perception discourage participation. Based on this anecdotal evidence, I developed a theory of change proposing that teaching readers to make small improvements during their normal reading would lower the perceived barrier to entry, reduce the likelihood of discouraging early conflicts, and help newcomers learn community norms gradually. I supplemented these observations by evaluating experimental results from the Wikimedia Foundation’s Newcomer Tasks features. Suggested tasks increased the proportion of account holders who made an edit and substantially reduced revert rates, although evidence for improved long-term retention remained limited. Based on this synthesis, I recommended treating small, low-risk contributions as a primary entry point into the community.

Speech in non-urban California

Most research into how Californians talk focuses on residents of large, coastal cities. In my PhD program I researched how Californians outside the urban centers talk, going to towns across the state and doing hour-long oral history interviews with life-long residents of these communities. With Rob Podesva (Stanford University) I found that speakers in these communities used different parts of the “California” accent to signal aspects of their personality and values, and that this differed across the state based on differences in these communities. For example, the Latinx communities in Bakersfield and Merced are long-established and a large proportion of the local populations, but while they are similar, not all Latinx speakers or communities are the same. I found evidence that speakers in these communities differ in how they pronounce the “a” sound in “trap” and “tram”, and these differences relate to how Latinx identity is performed in these different communities. This and other findings demonstrate that our understanding of language in California is impoverished when we do not consider the racial and geographic diversity of communities in the state. Read the qualifying paper.

History of a family of languages in Papua New Guinea

Kate Lindsey (Boston University) and colleagues have recently documented a family of languages in southern Papua New Guinea along the Pahoturi River. These languages share a common ancestor, but we don’t know where that common ancestor was spoken or how it was related to other languages in the area. One hypothesis is that speakers of these languages moved North from Australia. Another hypothesis is that speakers moved South from other parts of Papua New Guinea. To find evidence for these historical migrations, I have been reconstructing the common ancestor of these languages and evaluating which neighboring languages share traits with the reconstructed language. I have focused on a series of sounds, /kw/ and /gw/, which linguists call labialized velars. These sounds are found in some neighbors of the Pahoturi River languages, and some contemporary Pahoturi River languages have these sounds too. Were these sounds borrowed into Pahoturi languages from neighbors, or were these sounds in their common ancestor? If the sounds were in the common ancestor of the Pahoturi River languages, then it suggests that these languages are related to their Papuan neighbors rather than their Australian neighbors. By comparing the contemporary Pahoturi River languages, I showed that we can build a phylogenetic tree relating all the contemporary languages to a common ancestor through a series of language changes that happened in some communities but not others. Based on this evidence and previous work, I argue in favor of the Papuan hypothesis over the Australian hypothesis. Read the working paper.