Skip to menu Skip to content Skip to footer
Dr Martin Schweinberger
Dr

Martin Schweinberger

Email: 
Phone: 
+61 7 336 56892

Overview

Background

I'm a Senior Lecturer in Applied Linguistics at the University of Queensland, where I use computational methods and large text and speech corpora to study how English is actually used — in everyday conversation, across social groups, and by learners.

My research focuses on features that traditional grammars overlook but that shape real communication: discourse markers (like, you know), general extenders, terms of address, vulgarity, and adjective amplification. I also work in corpus phonetics, particularly vowel production and voice-onset timing in first- and second-language English speakers.

Alongside my research, I hold several roles focused on language research infrastructure:

  • Director of Research, School of Languages and Cultures, UQ
  • Director of the Language Technology and Data Analysis Laboratory (LADAL), a free open-access hub for language data science (https://ladal.edu.au/)
  • Chief Investigator on the Language Data Commons of Australia (LDaCA), a national research infrastructure project (https://www.ldaca.edu.au/)

I'm also an advocate for reproducibility and open science in the humanities and social sciences, and I develop tools and workflows that make text and speech analysis more transparent and rigorous.

Research areas: applied linguistics, corpus linguistics, computational linguistics, sociolinguistics, language variation and change, learner corpus research, corpus pragmatics, reproducibility, digital humanities.

Availability

Dr Martin Schweinberger is:
Available for supervision
Media expert

Qualifications

  • Doctor of Philosophy, Universität Hamburg

Research interests

  • Vulgarity and Swearing

    I investigate how swear words and taboo language are used in everyday speech and online discourse. Contrary to popular belief, vulgar language follows systematic social and linguistic rules. My research uncovers how such expressions function in communication and what they reveal about speakers’ identities, emotions, and group memberships.

  • Discourse Markers and Filler Words

    I study words like like, you know, and well—terms often dismissed as meaningless. Using computational analysis, I show how these elements structure conversations and convey nuanced meanings. My work demonstrates that such "filler" words play important roles in signaling attitudes, managing interactions, and guiding listener expectations.

  • Open Science and Research Transparency

    I actively promote reproducible, open research practices in the humanities and social sciences. I provide practical training and resources to help language researchers adopt transparent workflows. My advocacy supports greater academic rigor and long-term trust in empirical research.

  • Text Analytics and Computational Linguistics

    I apply computational methods—like machine learning and statistical modelling—to large corpora to uncover hidden linguistic patterns. These tools help quantify language use in a way that supports replicable, empirical research. My work is at the intersection of computer science and linguistics, making it especially relevant in the digital age.

  • Digital Infrastructure and Research Tools

    As Director of LADAL and a lead in LDaCA, I am building accessible digital platforms that support large-scale language analysis. These initiatives democratize access to language data and computational tools for researchers, students, and educators alike. My infrastructure work enhances the capacity for advanced language research in Australia and beyond.

  • Language Variation and Change

    I explore how language evolves over time and across different social settings. By analyzing large-scale linguistic datasets, I identify subtle patterns of variation in how people speak, particularly in informal and digital contexts. This research helps reveal how social norms and technology influence the way we communicate.

  • Learner Language and Second Language Acquisition

    I analyze how learners of English produce sounds, manage fluency, and develop pronunciation over time. This includes examining features like vowel quality, voice-onset time, pauses, and accent intelligibility. By comparing learner and native speaker data, my research informs language teaching and helps improve learner outcomes.

  • Corpus Phonetics

    I use corpus-based methods to investigate the phonetic characteristics of spoken language, including pronunciation patterns among both native and non-native speakers. I focus on measurable acoustic features such as vowel production and timing cues. This approach allows for the large-scale, data-driven analysis of speech in real-life settings.

Research impacts

Public reach through media

My research on vulgarity and variation in World Englishes has generated over 96 media reports internationally, with an estimated combined reach exceeding 205 million. In 2025, coverage included national television (Channel 10 News, The Project), ABC Radio interviews across Australia, and international press including The Guardian, CNN, Der Spiegel, Deutsche Welle, Popular Science, and Yahoo News. My co-authored article in The Conversation"201 ways to say 'fuck': what 1.7 billion words of online text shows about how the world swears" (with Kate Burridge) — prompted follow-up coverage across more than 20 outlets in Australia, the UK, Europe, and North America.

Building Australia's language data infrastructure

As Chief Investigator and Steering Committee member on the Language Data Commons of Australia (LDaCA), I contribute to building Australia's first national infrastructure for language data — enabling researchers, cultural institutions, and Indigenous communities to preserve, access, and analyse language collections that would otherwise remain fragmented or inaccessible.

Free global training in language data science

As founder and Director of the Language Technology and Data Analysis Laboratory (LADAL), I have built one of the most widely used open-access hubs for language data science training. Since 2021, LADAL has reached over 500,000 users worldwide, providing free, reproducible resources that lower the barrier to computational methods for students, researchers, and practitioners across the humanities and social sciences.

Shaping the field

I serve as Vice-President Profession of the International Society for the Linguistics of English (ISLE), board member of ICAME, Associate Editor of the Australian Review of Applied Linguistics, and book series editor for Bloomsbury's Language, Data Science and Digital Humanities.

Works

Search Professor Martin Schweinberger’s works on UQ eSpace

109 works between 2008 and 2026

41 - 60 of 109 works

2023

Other Outputs

Data in the humanities with text analytics

Schweinberger, Martin and Hames, Sam (2023). Data in the humanities with text analytics. Brisbane, QLD, Australia: The University of Queensland.

Data in the humanities with text analytics

2023

Conference Publication

A corpus-based acoustic analysis of vowel production by L1-Chinese learners and native speakers of English

Schweinberger, Martin and Yin, Rui (2023). A corpus-based acoustic analysis of vowel production by L1-Chinese learners and native speakers of English. 44th Meeting of the International Computer Archive of Modern and Medieval English (ICAME44), Vanderbijlpark, South Africa, 17 - 21 May 2023. Brisbane, QLD, Australia: University of Queensland.

A corpus-based acoustic analysis of vowel production by L1-Chinese learners and native speakers of English

2023

Other Outputs

Text Analytics with Language Technology and Data Analysis Laboratory Resources: an introduction to free, open-source interactive resources for linguists

Schweinberger, Martin (2023). Text Analytics with Language Technology and Data Analysis Laboratory Resources: an introduction to free, open-source interactive resources for linguists. Hamburg, Germany: University of Hamburg.

Text Analytics with Language Technology and Data Analysis Laboratory Resources: an introduction to free, open-source interactive resources for linguists

2023

Other Outputs

An introduction to conditional inference trees in R

Schweinberger, Martin (2023). An introduction to conditional inference trees in R. Bonn, Germany: Rheinische Friedrich-Wilhelms-University.

An introduction to conditional inference trees in R

2022

Conference Publication

A corpus-based computational analysis of high-front and -back vowel production of L1-Japanese learners of English and L1-English speakers

Schweinberger, Martin and Komiya, Yuki (2022). A corpus-based computational analysis of high-front and -back vowel production of L1-Japanese learners of English and L1-English speakers. Australasian International Conference on Speech Science and Technology, Canberra, ACT, Australia, 13 - 16 December 2022. Canberra, ACT, Australia: Australasian Speech Science and Technology Association.

A corpus-based computational analysis of high-front and -back vowel production of L1-Japanese learners of English and L1-English speakers

2022

Conference Publication

Exploring powerful tools to ensure robust and reproducible results in corpus linguistics

Schweinberger, Martin, Flanaghan, Joseph and Schneider, Gerold (2022). Exploring powerful tools to ensure robust and reproducible results in corpus linguistics. 43rd Meeting of the International Computer Archive of Modern and Medieval English (ICAME43), Cambridge, United Kingdom, 18 - 21 August 2021. Cambridge, United Kingdom: TU Dortmund University.

Exploring powerful tools to ensure robust and reproducible results in corpus linguistics

2022

Conference Publication

Research trends in corpus linguistics: a bibliometric analysis of two decades of Scopus-indexed corpus linguistics research in arts and humanities

Crosthwaite, P., Ningrum, S. and Schweinberger, M. (2022). Research trends in corpus linguistics: a bibliometric analysis of two decades of Scopus-indexed corpus linguistics research in arts and humanities. ICAME43, Cambridge, United Kingdom, 27-30 July 2022.

Research trends in corpus linguistics: a bibliometric analysis of two decades of Scopus-indexed corpus linguistics research in arts and humanities

2022

Other Outputs

From tables to forests – working with tables and tree-based models

Schweinberger, Martin (2022). From tables to forests – working with tables and tree-based models. Tromsø, Norway: The Arctic University of Norway.

From tables to forests – working with tables and tree-based models

2022

Other Outputs

Introduction to Power Analysis with R

Schweinberger, Martin (2022). Introduction to Power Analysis with R. Tromsø, Norway: The Arctic University of Norway.

Introduction to Power Analysis with R

2022

Other Outputs

Introduction to data visualization with R

Schweinberger, Martin (2022). Introduction to data visualization with R. Tromsø, Norway: The Arctic University of Norway.

Introduction to data visualization with R

2022

Book Chapter

Absolutely fantastic and really really good: language variation and change in Irish English

Schweinberger, Martin (2022). Absolutely fantastic and really really good: language variation and change in Irish English. Expanding the landscapes of Irish English research. (pp. 129-145) edited by Stephen Lucek and Carolina P. Amador-Moreno. New York, United States: Routledge. doi: 10.4324/9781003025078-7

Absolutely fantastic and really really good: language variation and change in Irish English

2021

Journal Article

Ongoing change in the Australian English amplifier system

Schweinberger, Martin (2021). Ongoing change in the Australian English amplifier system. Australian Journal of Linguistics, 41 (2), 166-194. doi: 10.1080/07268602.2021.1931028

Ongoing change in the Australian English amplifier system

2021

Journal Article

Training disciplinary genre awareness through blended learning: an exploration into EAP students’ perceptions of online annotation of genres across disciplines

Crosthwaite, Peter, Sanhueza, Alicia Gazmuri and Schweinberger, Martin (2021). Training disciplinary genre awareness through blended learning: an exploration into EAP students’ perceptions of online annotation of genres across disciplines. Journal of English for Academic Purposes, 53 101021, 1-16. doi: 10.1016/j.jeap.2021.101021

Training disciplinary genre awareness through blended learning: an exploration into EAP students’ perceptions of online annotation of genres across disciplines

2021

Journal Article

Which word gets the nuclear stress in a turn-at-talk?

Ruhlemann, Christoph and Schweinberger, Martin (2021). Which word gets the nuclear stress in a turn-at-talk?. Journal of Pragmatics, 178, 426-439. doi: 10.1016/j.pragma.2021.04.005

Which word gets the nuclear stress in a turn-at-talk?

2021

Journal Article

Voices from the periphery: perceptions of Indonesian primary vs secondary pre-service teacher trainees about corpora and data-driven learning in the L2 English classroom

Crosthwaite, Peter, Luciana and Schweinberger, Martin (2021). Voices from the periphery: perceptions of Indonesian primary vs secondary pre-service teacher trainees about corpora and data-driven learning in the L2 English classroom. Applied Corpus Linguistics, 1 (1) 100003, 1-13. doi: 10.1016/j.acorp.2021.100003

Voices from the periphery: perceptions of Indonesian primary vs secondary pre-service teacher trainees about corpora and data-driven learning in the L2 English classroom

2021

Other Outputs

Tree-based models in R

Schweinberger, Martin (2021). Tree-based models in R. Brisbane, QLD, Australia: The University of Queensland, School of Languages and Cultures.

Tree-based models in R

2021

Book Chapter

On the waning of forms – a corpus-based analysis of decline and loss in adjective amplification

Schweinberger, Martin (2021). On the waning of forms – a corpus-based analysis of decline and loss in adjective amplification. Lost in change: causes and processes in the loss of grammatical elements and constructions. (pp. 235-260) edited by Svenja Kranich and Tine Breban . Amsterdam, Netherlands: John Benjamins Publishing Company. doi: 10.1075/slcs.218.08sch

On the waning of forms – a corpus-based analysis of decline and loss in adjective amplification

2021

Journal Article

Analyzing Historical Changes in the Irish English Amplifier System

Schweinberger, M. (2021). Analyzing Historical Changes in the Irish English Amplifier System. Anglistik, 32 (1), 139-158. doi: 10.33675/angl/2021/1/11

Analyzing Historical Changes in the Irish English Amplifier System

2021

Other Outputs

Fixed- and mixed-effects regression models in R

Schweinberger, Martin (2021). Fixed- and mixed-effects regression models in R. Brisbane, QLD, Australia: The University of Queensland, School of Languages and Cultures.

Fixed- and mixed-effects regression models in R

2021

Journal Article

Analysing discourse around COVID-19 in the Australian Twittersphere: a real-time corpus-based analysis

Schweinberger, Martin, Haugh, Michael and Hames, Sam (2021). Analysing discourse around COVID-19 in the Australian Twittersphere: a real-time corpus-based analysis. Big Data and Society, 8 (1) 20539517211021437, 205395172110214. doi: 10.1177/20539517211021437

Analysing discourse around COVID-19 in the Australian Twittersphere: a real-time corpus-based analysis

Funding

Current funding

  • 2024 - 2028
    Language Data Commons of Australia (LDaCA-RDC)
    Australian Research Data Commons Limited
    Open grant

Past funding

  • 2021 - 2024
    Language Data Commons of Australia HASS RDC (LDaCA-RDC)
    ARDC - Australian Data Partnerships
    Open grant

Supervision

Availability

Dr Martin Schweinberger is:
Available for supervision

Looking for a supervisor? Read our advice on how to choose a supervisor.

Supervision history

Current supervision

  • Doctor Philosophy

    A multifactorial study of morpho-syntactic errors across different L1 backgrounds and language proficiency levels

    Principal Advisor

    Other advisors: Associate Professor Peter Crosthwaite

  • Doctor Philosophy

    Corpus-based investigation of three-minute thesis presentations: Register perspective

    Associate Advisor

    Other advisors: Associate Professor Peter Crosthwaite

  • Doctor Philosophy

    The Relationship Between Writing Tasks and Second Language Writers¿ Use of Metadiscourse

    Associate Advisor

    Other advisors: Associate Professor Peter Crosthwaite

  • Doctor Philosophy

    Integrating Artificial Intelligence and Machine Learning in TESOL: A Study on Personalised Learning and Impact on Student Engagement and Motivation in A Rural Indonesian University

    Associate Advisor

    Other advisors: Associate Professor Peter Crosthwaite

  • Doctor Philosophy

    Enhancing Lexical Resources for Argumentative Essay Writing through Corpus Integration

    Associate Advisor

    Other advisors: Associate Professor Peter Crosthwaite

Completed supervision

Media

Enquiries

Contact Dr Martin Schweinberger directly for media enquiries about their areas of expertise.

Need help?

For help with finding experts, story ideas and media enquiries, contact our Media team:

communications@uq.edu.au