Developing a Data Dashboard for
Discipline-Specific Text Analysis

NATESOL - Technology for 21st Century English Language Learning and Teaching

Charles Lam

Language Centre, University of Leeds

16 May 2026

Teaching Context: Qualitative Research in Biomedical Science and Healthcare

The Pedagogical Gap

  • Students in Medical and Healthcare typically receive little training

  • Qualitative studies are increasingly important

  • Existing resources are primarily for more “canonical” genres, e.g. writing guides: Divan (2009), Knisely (2017)

Number of Qualitative Studies (1936-2024)

Potentials for Data-Driven Learning?

  • Data-Driven Learning (DDL) “is an approach to language teaching and learning based on the tools and techniques of corpus linguistics to enhance learners’ understanding of how language is used across linguistic genres.” (Perez-Paredes & Boulton, 2025)
  • DDL fosters a lexico-grammatical awareness through exposure to attested, naturally occurring language rather than simplified textbook examples (Flowerdew, 2015; Johns, 1991)

  • STEM students tend to be receptive to empirical, data-driven evidence.

  • Authentic research articles from familiar databases (PubMed, PLoS) in ESAP makes the relevance immediately apparent (Hamp-Lyons, 2011)

Indirect + Direct DDL

Indirect DDL

  • Teacher curates corpus findings and presents them to students as pre-selected materials (Gabrielatos, 2005)
  • E.g. Teacher extracts frequent collocations of “participants” from the corpus for the handout
  • Indirect DDL is the more widely adopted mode in practice (Boulton & Forti, 2025)

Direct DDL

  • Students may interact with data / tools themselves.
  • Challenge: STEM students’ limited experience with language corpus. General-purpose interfaces can feel opaque without training (Kilgarriff et al., 2014)
  • The tools therefore need to be more explicit, scaffolded, and discipline-appropriate

Self-paced learning for students and novice writers

A Master’s student is drafting their qualitative methods chapter over the summer, with limited access to supervisory feedback or writing centre support. How do they know what is conventional?

  • Digital tool to scaffold independent learning:
    • Empirical norms (e.g., typical section length, token counts) so students can self-assess “how much should I write?”
    • Searchable concordance lines from published methods sections, enabling just-in-time lookup of disciplinary conventions
    • Common topics to explore similar articles

Seeing “what real articles look like” reduces anxiety about qualitative writing, particularly when students can compare patterns across multiple sources

This Study: Data Dashboard

Research Context & Goals

  • ESAP Approach at University of Leeds
  • Teaching context: Master Programmes in biology with various orientations (Conservation; Biomedical; Microbiology)
  • Small cohort for qualitative but persistent
  • Very high stake (dissertation research) and students can be intimidated by the unfamiliar approach

The Corpus

Source

  • Published research articles
  • Search phrase “qualitative study” in TITLE or ABSTRACT
  • Methods section only
  • Data already come in XML

PubMed Central

  • n=150 articles; 111,043 tokens
  • 740.29 tokens per article
  • The full dataset contains 15,159 articles
  • Sayers (2022) E‑utilities for PubMed Central / NCBI

PLoS (Public Library of Science)

Live Demo

Live Demo

Qualitative Research Dashboard

https://qual-dashboard.streamlit.app/



Click “Use PubMed Dataset” to start.

Key Features of the Dashboard

  • Corpus Analysis
  • Keyword Search
  • LDA Key Themes
  • Geographical Analysis
  • Move-Step Analysis (work in progress)

Corpus Analysis

  • Displays basic statistics
  • “How much do I need to write?”

LDA Key Themes

Latent Dirichlet Alloation

LDA (Blei et al. 2003) is a topic modelling technique that automatically groups documents by their underlying themes.

This topic modelling technique has already been adopted in corpus linguistics studies (Jaworska & Nanda, 2018; Huang & Jiang, 2026).

LDA Key Themes

Identified topics in the LDA results based on 15,159 full articles:

Geographical Analysis

Some topics are correlated to locations

  • GP in the UK
  • “Women’s health” in more industrialized vs. developing locations
  • automated Named-Entity Recognition (NER)
Map view

Map view (PLoS dataset)

Move-Step Analysis

  • Makeing genre features visible for students with little exposure to linguistics (Swales, 1990)
  • Currently with the DRaC (Demonstrating Rigour and Credibility) model by Cotos et al. (2017)
  • Towards operationalizing genre analysis?
  • Challenge: Difficult to display/read more articles

Move-Step Analysis


Also:
Ongoing work on more generalized python package for MSA visualization (comments welcome!)

Implications & Takeaways

AI for Fast Prototyping

  • My experience: the first prototype can be made within minutes
  • Dashboards / webapps are refined iteratively (build one tab at a time!)
  • Delegating to AI does leave the time to think about the design
  • “Cognitive off-loading” → less attached to design features and lowered time cost for changing

Empowering Language Experts

  • Language expertise guiding the tech / tool
  • AI handles the technical implementation
  • “Technology for 21st Century”
    → Removing tech barrier to create customized tools

Limitations & Future Directions

  • Plans for scaling or replication

  • Custom data set function (challenge with compute resources and storage)

  • Displaying Move-Step Analysis (currently 10 for demo)

    • Statistics of steps: raw count, percentage, step-n-grams (Lam & Nnamoko, 2024)
  • User training + training on replication

Replicating and Tailoring for your Context

The “Tech Stack”

These are just new websites that you can visit and start playing with.
You don’t need all of them!
Vibe coding is simpler than you think!

The “vibe coding” workflow

  1. Come up with an idea and list out the input and ouput
    Or just identify existing websites that you want to mimic
  2. Visit one of the AI-assisted coding platforms (see previous slide)
  3. Keep chatting! (aka: test, refine and iterate)
  4. Deploy!

No deep programming knowledge required.

Conclusion

  1. Main finding: DDL for teacher and students designed for lexical and genre conventions
  2. For EAP practitioners: use AI for technology!
  3. Aligning with calls in the literature for DDL tools that function beyond formal instruction (Perez-Paredes & Boulton, 2025; O’Keeffe, 2021)

Thank You

Questions and comments welcome!

Other use cases?

Other corpora?


c.lam@leeds.ac.uk

https://charles-lam.net/presentations/NATESOL2026/

References

Blei, D. M., Ng, A. Y., & Jordan, M. I. (2003). Latent dirichlet allocation. Journal of machine Learning research, 3, 993-1022.

Boulton, A., & Cobb, T. (2017). Corpus use in language learning: A meta‐analysis. Language learning, 67(2), 348-393.

Boulton, A., & Forti, L. (2025). Corpus linguistics and data-driven learning. In International Encyclopedia of Language and Linguistics (pp. 1-8). Elsevier.

Cotos, E., Huffman, S., & Link, S. (2017). A move/step model for methods sections: Demonstrating rigour and credibility. English for Specific Purposes, 46, 90-106.

Divan, A. (2009). Communication Skills for the Biosciences. Oxford University Press.

Flowerdew, L. (2015). Data-driven learning and language learning theories. In Flowerdew, L., Leńko-Szymańska, A., & Boulton, A. (eds). Multiple affordances of language corpora for data-driven learning, pp. 15-36.

Gabrielatos, C. (2005). Corpora and language teaching: Just a fling or wedding bells. TESL-EJ, 8,1–35.

Hamp-Lyons, L. (2011). English for academic purposes. In Hinkel, E. (ed) Handbook of research in second language teaching and learning, pp. 89-105.

Huang, Z. & Jiang, Z. (2026). Text mining of syntactic complexity in L2 writing: an LDA topic modeling approach. International Review of Applied Linguistics in Language Teaching, 64(1), 523-548. https://doi.org/10.1515/iral-2024-0132

Jaworska, S, Anupam Nanda, A. (2018) Doing Well by Talking Good: A Topic Modelling-Assisted Discourse Study of Corporate Social Responsibility, Applied Linguistics, 39(3), 373–399, https://doi.org/10.1093/applin/amw014

Johns, T. (1991). Should you be persuaded: Two samples of data-driven learning materials. ELR Journal 4, 1-16.

Kilgarriff, A., Baisa, V., Bušta, J., Jakubíček, M., Kovář, V., Michelfeit, J., … & Suchomel, V. (2014). The sketch engine: ten years on. Lexicography ASIALEX 1, 7-36

Knisely, K. (2017). A student handbook for writing in biology. Macmillan.

Lam, C., & Nnamoko, N. (2024). Quantitative metrics to the CARS model in academic discourse in biology introductions. In Proceedings of the 5th Workshop on Computational Approaches to Discourse (CODI 2024) (pp. 71-77). Retrieved from https://aclanthology.org/2024.codi-1.7.pdf

O’Keeffe, A. (2021). Data-driven learning–a call for a broader research gaze. Language Teaching, 54(2), 259-272.

Pérez-Paredes, P. and Boulton, A. (2025). Data-driven Learning in and out of the Language Classroom. Cambridge University Press. https://doi.org/10.1017/9781009511384

Sayers, E. (2022, November 17). A general introduction to the E‑utilities. In Entrez Programming Utilities Help [Internet]. National Center for Biotechnology Information (US). Retrieved from https://www.ncbi.nlm.nih.gov/books/NBK25497/#chapter2.The_Nine_Eutilities_in_Brief

Swales, J. M. (1990). Genre analysis. Cambridge University Press.