Machine Learning in the Humanities and Social Sciences: Why It Matters and How to Use It
Key Takeaways
- Machine learning is rapidly expanding its role as a methodological complement in the humanities and social sciences — handling unstructured data, bridging explanation and prediction, exploring heterogeneous effects, and supporting LLM-assisted workflows.
- Specific methods such as STM, word embedding, causal forest, computer vision, and LLMs have demonstrated their value through concrete case studies that transform text, image, and spatial data into meaningful social science variables.
- When applying machine learning, systematic validation of measurement validity, sampling bias, normative appropriateness of prediction targets, reproducibility, and ethics must accompany every step.
1. Introduction
In the humanities and social sciences, machine learning (ML) serves as a methodological complement — one that measures large-scale unstructured data, tests the limits of predictability, and uncovers heterogeneous effects and hidden patterns.
The case for machine learning in the humanities and social sciences rests on several foundations.
- First, data that were once inaccessible through traditional surveys, interviews, or statistical tables — digital text, images, audio, networks, and behavioral logs — have become central research materials.
- Second, abstract concepts such as discourse, sentiment, cultural meaning, urban environment, social status, and political orientation can now be translated into measurable variables from large-scale data.
- Third, “how predictable is a social phenomenon?” has become as important a research question as “why does it occur?”
- Fourth, researchers can now explore heterogeneous effects across subgroups and contexts, rather than being limited to average effects. [Molina, M., & Garip, F (2019)]
- Finally, large language models (LLMs) are increasingly used as auxiliary tools for text coding, hypothesis generation, experiment design, simulation, and literature search.
At the same time, there is growing discussion about the need to validate measurement validity, sampling bias, reproducibility, interpretability, fairness, and ethics whenever machine learning is applied. This is especially important because social science data involve people, institutions, and power relations — which means questions like “what is being predicted?” and “what social construct does the prediction proxy?” matter far more than simple accuracy metrics. [Lazer, D. M., Pentland, A., Wa… (2020)]