Notes from the DHIP-Summer University in Paris | Doing Migration History with Digital Methods.

This essay presents a shortened, reworked version of the keynote lecture given at the summer university ‘Doing Migration History with Digital Methods’, held at the German Historical Institute Paris (DHIP / Institut historique allemand) and organized by Dr. Mareike König, deputy director of the DHIP, together with the Luxembourg Centre for Contemporary and Digital History (C²DH, University of Luxembourg), 22–26 June 2026. The keynote was delivered on the evening of Thursday 25 June and was chaired by Prof. Dr. Machteld Venken (C²DH, University of Luxembourg), one of the summer university’s organizers.


Conference impressions

The Summer-University week – in which I participated from the 24th of June – moved through the digital methods now reshaping migration history. After Lorella Viola’s opening keynote on mapping migrant worlds, the sessions ranged across geographic information systems and spatial analysis, optical-character and handwritten-text recognition, the modeling of migration through textual data, datafication and migration databases, and digital network analysis, with hands-on workshops from QGIS mapping to LLM-assisted workflows. One question ran through all of it: what a digital approach opens, and what it closes, in matters of bias, interpretation, and the limits of the method. The contributions were, for the most part, doctoral projects, each presented to the summer university, commented on by an expert, and discussed by the group.

Participants of the DHIT-Summer-University 2026 (Photo: Rass)

On Friday, I chaired the first session on digital network analysis, where Alexandre Binoux (Université Paris 1 Panthéon-Sorbonne / École française de Rome) read the eighteenth-century Greek colony of Corsica by turning parish registers and letters into networks, and Jorit Jens Hopp (Ludwig Maximilian University of Munich, ERC project T-MIGRANTS) traced the professional networks of German-language theatre people in the nineteenth-century Habsburg Empire.

Two further papers given in Thursday’s sessions stayed with me in particular: Timur Mitrofanov (University of Heidelberg / Herder Institute for Historical Research on East Central Europe) read the English-language press of the Latvian diaspora, some 7,270 articles across twenty-five titles, through topic modelling and named-entity recognition; and Valentin Rhodius (Université de Caen Normandie) built a database of the maritime crossings the International Refugee Organization organised between 1946 and 1952, more than 1,200 voyages that carried close to 700,000 of post-war Europe’s displaced overseas. Both connect directly to research lines pursued in our own NGHM working group (Neueste Geschichte und Historische Migrationsforschung) in Osnabrück.

The essay that follows reworks my own keynote, delivered on the Thursday evening: shortened, with the questions raised in the discussion folded into the argument and gathered again at the close.1


How Migration Became Data, How Data Makes Meaning, and How Reflexive Migration Research Intervenes

by Christoph Rass

On 13 February 1905, Tullio and Theresa Beltrami, who had come from Italy, moved into a flat in a workers’ district of Osnabrück.2 They stayed. Two daughters were born in the city, in 1909 and 1913: a second generation before the First World War. And yet, under the principle of ius sanguinis, those daughters were Italian, not German, born in Osnabrück and made foreign by descent. From 1916, the family lived as nationals of a state at war with the German Reich, classified as ‘enemy aliens’ on a street where they had already belonged for a decade.

This family is not to be found in a single coherent document. They left almost no trace of their own. They survive in the municipal foreigners’ registration index, a card file created in an administrative reform at the close of the 1920s, when the most advanced information technology of the day reached general administration. We reconstruct the Beltramis as a trajectory across that index: by reading handwritten entries and by record linkage, the matching of scattered cards into one family’s path.

This is the first thing worth noticing: the index did not record the Beltramis, it made them. So we read the “Ausländermeldekartei” twice: as a source for their lives, and as a source for the act that registered those lives. The friction between the two, between a biography and the grid that inscribes it, is what makes the file legible as an archive of its own categorical work.

Migration became data long before it became digital. Markus Krajewski has shown that the card index was already a data-processing machine; what our instruments add today is scale and speed, not the operation itself.3 This essay takes up that double reading. It traces how a categorical figure is produced, assigned, and contested; it asks what the methods of digital history open and what they close; and it locates the silences of the data infrastructure, naming a fourth one that is new in kind. One commitment runs through all of it, stated here against the temptation to read the diagnosis as a counsel of despair: the same operation that produces a silence can be turned against it, and the same digitization that deepens a silence can also help to break it. The argument is not a verdict on digital history. It is an attempt to keep the method usable and honest at scale.

A figure and its trajectory

Widen the focus from one family to the drawer that holds it. The Osnabrück foreigners’ index records about 83,000 arrival entries for more than 70,000 people between 1881 and 1980, with adequate data quality for analysis from the late 1930s onward.4 Across that period the same printed form runs through five different populations: the late Weimar administration that begins the file; the National Socialist state that turns it to the administration of forced labour; the British occupation that registers ‘displaced persons’; the Federal Republic that fills it, through the recruitment decades, with Spanish, Italian, Turkish, and Portuguese ‘guest workers’; and, after the recruitment ban of 1973, the children of family reunification, thousands of them born in Osnabrück as ‘foreigners’ by descent, as the Beltrami daughters had been six decades earlier. The people changed. So did the clerks, and even some of the forms. The grid was the constant.

Prof. Dr. Christoph Rass @ DHIP (CC-BY 4.0 DHIP)

This is the wager of the Collaborative Research Center 1604 ‘Production of Migration’ in Osnabrück, where I work as one of the principal investigators, and the analytical position applied throughout is the one we have developed there.5 Migration is not the cause that explains social processes; it is their contested product. Human mobility is a fact. ‘Migration’ is a construct, and our object is the production of that construct, including by digital tools and, increasingly, with them. The word ‘production’ is chosen over ‘construction’ on purpose: construction points to language and discourse, production points to apparatuses, factors, and products, to the registration office as much as to the speech acts performed there. Migration, in Durkheim’s phrase, is treated as a social fact, something a society makes and then meets as if it were given; and the same may hold for the work of production itself.

Take the one word that organizes the index across decades, from the wartime years into the rise and decline of the West German ‘economic miracle’: ‘guest worker’, ‘Gastarbeiter’. It looks like a description of a labor-market arrangement. It is a sediment. Max Weber introduced the term in a context remote from labor migration;6 the National Socialist regime then took it up in wartime propaganda, where it performed a semantic radicalization. It masked coercion as legitimate exchange, gave the forced-labor system a pseudo-legal veneer, and folded the ‘Ausländereinsatz’, the deployment of foreign labor, into a pan-European fiction. When the Federal Republic revived ‘Gastarbeiter’ in 1961 to replace ‘Fremdarbeiter’, a term seen as compromised by its National Socialist forced-labor associations, the older meanings carried over into the practices of recruitment and their registration routines. The word did not simply label the arrangement: it shaped it. ‘Gastarbeiter’ meant one thing above all: a temporary, segregated presence performing a social function, with no path to belonging.

Figures of migration do not stand still once produced. They acquire trajectories: they activate institutions, they address the people they name, and they are, at times, turned against their makers. What holds for the figures we inherit holds for the figures we make.

In the model my colleague Hajo Holst and I have been developing at CRC 1604, in the Center’s work on ‘figures of migration’, we see three forces at work, rarely in equal measure.7 Institutionalization is the stabilizing pull of the apparatus: by 1955, when the Federal Republic signed its recruitment accord with Italy, more than two dozen instruments of regulated labor migration were already in force across Europe, each stabilizing a particular migration regime through the category it deployed. Subjectivation is the pressure of categorical address: a child first entered on a parent’s card as a dependent, a ‘Gastarbeiterkind’, then issued a card of its own at a first job, is handed a position before it can contest it. Transformation is the appropriation that rewrites a figure against its origin: Wir sind keine Gäste, we are no guests. With that slogan, immigrants demanded representation, fought deportation, and claimed equal treatment for their children at school and equal wages at work. The children of the recruited raised it in the 1970s and 1980s; the ‘third generation’ that followed them carries it still. Those produced by the call for ‘guest workers’ took the label and turned it into a claim to precisely what it had been built to deny. We watch that claim materialize in the index itself, as the first of those children crossed from the foreigners’ file into the residents’ register as naturalized citizens.

None of this is unique to migration. Ian Hacking described how classifications of people loop: the classified notice the category assigned to them, respond to it, and at times shift it in turn, so that human kinds become moving targets.8 The categorization is usually easy to spot; the contestation we have to detect with care. A figure is more than a classification: it carries a narrative and an image, a recognizable shape that takes part in the order it appears only to describe. The discipline of working with such figures is to register that twist, what Geoffrey Bowker and Susan Leigh Star call ‘torque’, the way a life is wrenched as it is forced to fit a classification,9 rather than to reproduce it. The reach is wider than one word: the ‘displaced person’ of 1945, a category built by the international relief apparatus, came alive in the same way, and the mechanism extends to gender, race, class, and disability alike. What we call ‘migration’ here might be better caught, as Janine Dahinden proposes, by ‘migrantization’, the process by which people are made migrants in the first place.10

The ‘second generation’ returns at the threshold of the second part of this argument. The city-born children who the law still counted as foreign reappear in the 1970s as a statistical wave, exactly when West German migration studies begin to produce them as a figure. In the Osnabrück file, the number of registered minors peaks in 1973, at roughly 2,500 children; in 1974, the year after the recruitment ban, some 850 new children arrive in a single year, family reunification overrunning the political intention to halt and reverse migration; and the cohort reaching school age rises eightfold between 1955 and 1977. As their numbers grew, the city administration itself began to label these children ‘Gastarbeiterkinder’.11 The administrative file and the scholarly corpus emerging at the same time record the same child from two sides. That correspondence is the bridge to the digital methods themselves.

What digital methods open, and what they close

Every digital method brings a gain in what we can see at the price of something we must fix in place. Our own vocabulary has a word for the second half, Gestaltschließung: the pressure to resolve every ambiguity into a value, a category, a column. Three modes organize the digital history of migration, and each pays the same price in its own currency.

Mapping comes first because it is the most frequent of these operations. What a map opens seems real at once: geocode the Osnabrück addresses, and a segregated social topography appears at a glance. One address resolves into more than five thousand entries, overwhelmingly Spanish, a company-run ‘guest worker’ hostel on a site that had itself been a forced-labor camp between 1942 and 1945, rendered as a single point. What a map closes is the question of what, exactly, has been mapped. A migration map rarely shows where people were or how they moved; it shows where an administration recorded them, projected onto coordinates the source never had, with a precision that only looks exact. We plot what John Searle calls social facts, things that exist only because an institution recorded them, as if they were brute facts that had been there all along.12 Critical cartography has insisted since Brian Harley that a map is an argument, not a mirror, and Johanna Drucker reminds us that what we plot is never simply data but capta, something taken, not given. Henk van Houtum has shown how the innocent arrow on a migration map manufactures the very sense of invasion it claims to depict.13 The answer is not to stop mapping but to map provenance alongside position: to mark whether a point is a measurement, an administrative category, or an imagined place.

The second mode reads language, and here I write as one of the accused: in one strand of the Centre’s research we built a diachronic corpus of the discipline’s own writing on migration and education: roughly six thousand German-language texts from the 1960s to the 2010s, some forty-five million words in the cleaned working corpus.14 The corpus is not about migration research in the abstract but about migration research and the school system, about the making of the ‘migrant child’ as an object of knowledge. That research began in Germany at the moment those children appeared in the classrooms, a country unprepared for an arrival no one had forecast; it was, in a sense, the zero hour of German migration studies, called into being by a real problem at the school gates. We use the corpus to ask not whether the children labeled ‘Gastarbeiterkinder’ or ‘Ausländerkinder’ or ‘deficit children’ had language problems, but how the research discourse produced the ‘language problem’ as a fact. Run the literature through the pipeline, and two operations appear at once: the one that produces voice, and the one that produces silence. The concepts stay stable while the labels change. One finding is visible only at scale: the theorem of ‘double semilingualism’, the claim that a child raised in two languages masters neither, enters the German debate abruptly around 1979, imported from Scandinavia, and arrives bound to a single companion, the vocabulary of danger. The collocation tells us what the concept was for. It did not describe the children’s needs; it warned society about them. The deficit term almost never sits in a monolingual setting: no monolingual German child with a small vocabulary is ever called ‘semilingual’. The word does not measure competence. It marks belonging and reports the mark as a measurement. The corpus also carries a hopeful counter-finding: from the early 1990s the field begins to turn on itself, and the vocabulary of construction, ascription, and reflection rises steadily across the following decades.

The third mode models data; it drives the work of networks and databases. What it opens is relation and differentiation: the nationality-specific timing by which Spanish families crossed into a second generation earlier than Turkish or Portuguese ones, a pattern no card-by-card reading could find. What it closes is identity itself. Record linkage decides which two cards are one life; entity disambiguation decides which person a name denotes in the first place. These are not neutral operations: a wrong merge invents a biography that never existed, a wrong split hides one that did. José van Dijck calls datafication an ideology, and the power of a classification lies precisely in its invisibility.15 In the Osnabrück foreigner’s register, the post-war category ‘unclear nationality’ resolves the ambiguity of a whole displaced continent into a seemingly unambiguous value: unclear-Polish, 914; unclear-Yugoslav, 892; stateless, 506. The clerk, in each instance, had to choose a box, and so does the data schema. The remedy is not to refuse the model but to keep the choice legible: documented schemas, data statements, a source criticism of process-generated data carried unbroken into the database.

Migration of a silence?

Three modes, one finding each: every gain in differentiation is paid for in disambiguation; the more powerful and the easier the tool, the less visible the closure becomes. On the handwritten card the categorical decision stands in plain sight: nationality, status, occupation, written in fields open to challenge. In the relational dataset it moves into the schema, already less visible. In the language model it settles into weights and into the distribution of the training data, far harder to inspect. With the rise of model power, the site of categorical inscription migrates from the visible surface into the opaque interior. The bias has not necessarily grown. Its transparency has fallen.

That formulation drew the sharpest exchange of the discussion after my talk, and it is worth keeping open rather than resolving. One can object, with reason, that the model does not merely hide an old bias but adds a new layer of its own, so that the bias does grow. Both readings hold. Whether the added model layer means ‘more bias’ or ‘the same bias, buried deeper’ is a genuine conceptual question, and the reflexive point survives either answer: a distortion we can no longer see is harder to contest than one written in a clerk’s hand.

Michel-Rolph Trouillot located silencing at four moments in the production of history: the making of sources, the making of archives, the making of narratives, and the making of history in its retrospective significance.16 What I have not seen done, and what seems worth doing, is to ask where each of these moments now sits in a data infrastructure. The making of sources moves into data modelling: the decision about what can become a field, a value, a class at all. The making of archives doubles the older question of what enters the archive with a new one: what gets digitized, and how the metadata classify it; and, as Geoffrey Bowker and Susan Leigh Star insist, every category valorizes one perspective and silences another.17 The making of narratives becomes the logic of ranking, indexing, and retrieval, the searchable shadow that Lara Putnam warned would make whatever has been digitized disproportionately visible.18 To read a source critically now means reading the whole pipeline that delivered it, not only the page at its end.

A fourth moment may be new in kind. It sits in the machine generalization of the large language model, where the distribution of the training data prejudges what can count as significant before the first question is asked. This is not a break with everything that came before. It is the same categorical operation, migrating one step further along a chain that runs from the clerk’s hand to the card index, from the card index to the relational database, and from the database into the trained model, carrying its silences at every step. Recent accounts tend to reach for the language of rupture, a paradigm break from archival mediation to opaque algorithms, yet without the historical depth that Trouillot, and Ann Laura Stoler’s work on the colonial archive, supply.19 Rupture, however, may be the wrong figure. What we are watching could also be read as a migration. Bias is the symptom; silencing, distributed across the moments of production, is the concept. And a migration can be made to run the other way: that is the reversal promised at the outset.

It helps to separate two debates that are usually run together. One is what large language models could do as a technology; the other is what the (commercial) models actually available to us do now. None of them is trained on historical sources, which is why they do not really read our sources and why they return curious explanations: the training data does not fit the purpose to which we put it. And their bias is not a separate machine defect set against an unbiased world. It is the reproduction of the bias already in the world, the same bias the archive has always reproduced from the time that made it. What we are trained to do with a biased source is exactly what a reflexive practice must learn to do with a model.

The reverse operation

If silence enters at four moments, a reflexive practice can answer at each. Where sources are made, we can model the ambiguity instead of resolving it away, and let ‘unclear nationality’ stand on the record as the open question it was. Where archives are made, handwritten-text recognition makes legible, at a scale no reading room ever allowed, what had been buried in undeciphered hands; and entity recognition can surface even a name that appears only once, a trace no index ever caught, as Mrinalini Luthra and colleagues have shown for the colonial archive.20 Where narratives are made, record linkage returns individuated lives from the categorical mass. The reason we now know the Beltramis by name across three generations is that these methods gave them back a trajectory rather than a type.

That recovery rested on the same fragile decisions warned against above. Every link in the Beltrami trajectory is a disambiguation that could have gone wrong, and a different choice would have given a different family. The categories are my data as much as the clerk’s: I counted the second generation through two definitions, and the two definitions yielded two different cohorts and two different images. This is unsilencing, and it is real, and it is not innocent, because to recover a life as data can re-objectify it, as Jessica Marie Johnson warns: every entity type and every linkage rule is itself a categorical decision.21 An unsilencing that knows it is also a kind of inscription: that is the discipline worth arguing for. Not refusal, and not blind faith, but a method turned against its own silences with eyes open.

The constructive task that follows is a research program, not a hope. It begins by relocating the question of confidence. Character recognition may, before long, become a solvable problem: given a controlled historical lexicon and a pass of human verification, domain-trained models already read difficult administrative hands at high accuracy, even on densely structured forms.22

The harder question lies one step further on. It is not whether the machine reads the character but whether the answer it then generates can travel back to its data points and carry a graded measure of its own certainty. That is the confidence score worth arguing about, and it sharpens at scale: forty or eighty thousand cards can still be checked, but when an archive offers three million, verification crosses a threshold beyond which only statistics and probabilities remain. The records that make recognition hardest are built up in strata of stamps, corrections, and rewritten decisions. The ‘CM/1’ files — the application files displaced persons submitted to the International Refugee Organization for assistance, now held at the Arolsen Archives — are the object of a collaboration with the pattern-recognition group of Professor Gernot A. Fink at TU Dortmund.23

The Summer University program also asks us to build richer digital objects. Why, for example, should a dataset travel alone? A project on the history of migration conferences that we are currently designing can build, around a relational core, a layered corpus: the proceedings themselves, the contemporary press coverage, the scholarship of the time, and the scholarship of today. Mixed methods then move between the macro-view, which locates where close reading is worth the historian’s attention, and the deep dive that returns, always, to the document. And it asks us to build the instruments ourselves and to divide the labor, rather than wait for firms whose general-purpose models answer to actors who do not share our questions or our sources. Let one kind of model open the source, reading, extracting, and linking at scale, ideally as a source-bound knowledge graph in which every statement carries its provenance and a graded measure of certainty, so that the contradictions in the record are preserved rather than smoothed away.24 Let the interpretation stay with us, tied to the documents and to the people those documents register, never to a model’s fluent voice. Joanna Guldi calls the text-mining turn at once a revolution and a danger, and both verdicts are true; Torsten Hiltmann’s experiments point to the corrective: the model performs best exactly where a historian’s knowledge guides it.25 There is no shortcut to expertise.

None of this is, in the end, a translation of the old craft into a new medium. The critical historian cannot simply be carried across. The coming generation will not transition from analog roots, as an older one did slowly over thirty years, from a first relational database in the early 1990s; it will grow up in a world where the non-digital method is marginal, something seen only once in a proseminar. The virtues of the craft have to be preserved, but the method itself has to be reinvented over what will likely be a period of trial and error, much as the critical method we now take for granted was the work of generations. To refuse that task would be a betrayal of one’s students and our discipline. It also has an institutional and economic edge that the humanities tend to avoid: unless we move from playing on our laptops to becoming serious computational actors, with access to real processing time and the resources it requires, we will be reduced to consumers of tools built by others, and the open question will be whether historians remain involved in the production of history at all.

The summer university’s visit to the Musée de l’histoire de l’immigration, in a building raised for a colonial exhibition and now given over to representing migration and diversity, is the place to close. CRC 1604 runs a transfer project in such a space, ‘Reflexive Migration Research in the Museum’, with the documentation center DOMiD, which is building Germany’s first national museum of the ‘migration society’, not of ‘immigration’: here, too, the category matters.26 The task it sets is to intervene without reproducing the very figures whose production we have traced. Adrian Favell has pressed the hardest version of the difficulty: every analytical vocabulary tends to acquire the governance function it set out to expose, and ‘production’, our own word, is not exempt.27 The name of our Center is itself a wager, and a reflexive discipline keeps that wager visible rather than pretending it has none. The tension cannot be dissolved. It can only be navigated, in the open.

An excursion took participants to the French Immigration Museum (Photo: Rass)

So this essay returns, at last, to the card. The clerk’s hand, when filling in Tullio Beltrami’s data, did not record a foreigner who was there. It made one, as data, and the figures made through norms, registers, classification practices, datafication, academic and public discourse went on to organize hostels, schools, statistics, scholarship, and now the maps, the corpora, and the models we build to study it. The same operation runs forward into the present, where asylum procedures and borders are datafied in real time, the new infrastructure layered over the old paper without ever fully replacing it. Even a fully digitized future would not retire the oldest questions: what counts as an event worth recording, who records it, who keeps the record, who is trained to read it, and who assigns significance to the narratives that result. Silencing is structural, not a function of scarcity. The honest task is not to claim a vantage point outside the categories, which we do not have and never will, but to write the history of migration and diversity from inside the work that produces them. On the card, the decision stood in plain sight. The model is the newest clerk, and it writes from a hand we can no longer read. The bias has not grown; its transparency has fallen; and to keep reading that hand, from the inside and with open eyes, is the work that remains.


Notes from the discussion

The exchange after the lecture circled, more than once, around a single worry, voiced most directly by the chair: that the diagnosis seemed to leave no way out, and that it placed a heavier burden on digital history than the analogue craft ever had to carry. My reply, and the questions that followed, sharpened several points.

Is the bias growing, or only less visible? Machteld Venken proposed that the new models add a layer of distortion of their own, so that bias grows rather than merely loses transparency. I let the tension stand: both readings are defensible, and the reflexive task holds either way. What I added was a concrete reason the worry is not a machine fault to be patched but a historical condition to be read: as their supply of text runs short, the firms building these models have begun buying up antiquarian books, reference works, and old scholarship to feed them, so that whatever a model later returns has already passed through the same archive, and the same biases, we are trained to read against the grain.28 If that is where the data come from, one might as well pay the national archive a visit. Venken also recalled a wiki-based knowledge graph shown earlier in the week that already did much of what the argument asks for, binding provenance and a graded certainty to each entry; the constructive direction was not hypothetical but already on the table.

Prof. Dr. Machteld Venken & Prof. Dr. Mareike König during the concluding discussion of the DHIP Summer-University (Photo: Rass)

Can you hold your own standards in the text-mining? Yes, and I tried to say how: standard text-mining components assembled in Python, documented code, and primary data prepared so the research can be reproduced. The harder problem is not the method but the practicable standard, what we must submit alongside our results so that a public can hold us to account, a demand that multiplies as the data grow. Why it matters shows in the corpus itself: run the discipline’s own writing through the pipeline, and the operations that produce voice and the operations that produce silence appear side by side. When, in the 1970s, the Chicago School’s ‘second generation’ was carried onto a different society, a few contemporaries warned that the transfer mistranslated both the term and the two societies it moved between; they were not heard, and the corpus still preserves the objection it silenced. One tension I cannot yet resolve sits underneath all of this: ambiguity-preserving documentation, which keeps every uncertainty on the record, and efficient modeling, which resolves them to move at scale, pull against each other, and the semantic-wiki and the hard-data-modeling traditions still barely speak. We are, in this, a little like Goethe’s sorcerer’s apprentice. One reason we press the question at all is that so few others do: the field is either too fascinated or too repelled to press us, which makes it the more important that we press ourselves.

What about confidence and probability scores? Recognition itself is, I think, largely a solvable problem, and a model that refuses to guess can be a virtue: faced with a barely legible hand, a model that declines to supply a plausible but unverifiable reading is doing exactly the right thing. Highly curated, domain-trained models already read difficult historical hands: the Swedish National Archives’ ‘Swedish Lion’, trained on court records from the seventeenth to the nineteenth centuries, reads early-modern Swedish at scale, and the Osnabrück indices were processed to a high reported accuracy on structured forms, every card checked by hand.29 As the essay argues, the question that matters lies not in the character but in the confidence of the answer the model builds from it — and that question only sharpens as the archive grows.

Translation, or reinvention? A participant asked whether the task was to translate the critical historian into the digital or to reinvent the figure. The latter, I believe — for the reasons the essay gives. A follow-up that anyone working on a category like ‘unclear nationality’ should first study the diplomatic record that produced it opened the most constructive idea of the afternoon: that a dataset need not travel alone, but can be built as a layered object in which quantitative reach and close reading become two movements of one practice.

And when the source problem disappears? A final question imagined a future in which nearly everything is digitized and source scarcity no longer constrains us. The structural questions the essay names would remain untouched: these, not the supply of documents, are where silencing lives. The closing note was institutional rather than technical: unless historians become serious participants in computation, with the processing time and resources it demands, the surface systems of the future will write history on request, on a sunny afternoon in Paris, without us. That is the prospect worth refusing, and the reason to hold our own.

Note: This section is reconstructed from my notes and paraphrased throughout. All errors are mine.

References

  1. Summer university ‘Doing Migration History with Digital Methods’, German Historical Institute Paris (DHIP) with the Luxembourg Center for Contemporary and Digital History (C²DH, University of Luxembourg), 22–26 June 2026. Speaker names, affiliations, session order and paper titles follow the conference program; the summaries draw on the speakers’ submitted papers. Sessions referenced: ‘Exploring Migration Using Digital Network Analysis I’ (chair: C. Rass; Binoux, Hopp); ‘Modelling and Analysing Migration through Textual Data’ (Mitrofanov); ‘Datafication and Migration Databases’ (Rhodius).↩︎
  2. The Beltrami case and the figures drawn from it derive from the Osnabrück foreigners’ registration index (Ausländermeldekartei, Niedersächsisches Landesarchiv Osnabrück, Dep 3 c Akz. 2019/83) as reconstructed in the author’s research; biographical details follow the verified keynote text. On the index as a source: Sebastian Bondzio, Linda Ennen-Lange, Lukas Hennies, Sebastian Huhn, Max Pochadt, Christoph Rass and Janine Wasmuth, ‘Die Osnabrücker Ausländermeldekartei 1930–1980. Potenziale als Quelle der Stadt- und Migrationsgeschichte’, Osnabrücker Mitteilungen 126 (2021): 137–193; Janine Wasmuth and Christoph Rass, ‘Die Osnabrücker Ausländermeldekartei’, Archiv-Nachrichten Niedersachsen 27 (2023).↩︎
  3. Markus Krajewski, Paper Machines: About Cards and Catalogs, 1548–1929, trans. Peter Krapp (MIT Press, 2011).↩︎
  4. Aggregate figures for the Osnabrück index (c. 83,000 arrival entries, more than 70,000 persons, 1881–1980).↩︎
  5. DFG Collaborative Research Centre (SFB) 1604 ‘Production of Migration’ / ‘Produktion von Migration’, Universität Osnabrück.↩︎
  6. That Max Weber introduced the term ‘Gastarbeiter’ is a result of the author’s research on the genealogy of the figure: Christoph Rass, ‘Gastarbeiter’ – ‘Guest Worker’. Translating a Keyword in Migration Politics (IMIS Working Paper 17, Osnabrück 2023); and Christoph Rass, ‘Gastarbeiter’ in the Nazi Press (1933–1945). Formative Years of a Key Term in Migration Politics (IMIS Working Paper 22, Osnabrück 2026). Weber’s own usage: Hinduismus und Buddhismus (Gesammelte Aufsätze zur Religionssoziologie II, Mohr/Siebeck 1921).↩︎
  7. The three-forces model (institutionalization, subjectivation, transformation), developed by the author with Hajo Holst: Christoph Rass and Hajo Holst, ‘When Categories Come Alive: Theorizing Generative Dynamics in “Figures of Migration”’, in Production of Migration, vol. I: The Concept of the Production of Migration, ed. Andreas Pott et al. (Cham: Springer, forthcoming).↩︎
  8. Ian Hacking, ‘Making Up People’ (1986); ‘Kinds of People: Moving Targets’ (Proceedings of the British Academy 151, 2007, 285–318); the ‘woman refugee’ example is in The Social Construction of What? (Harvard UP, 1999).↩︎
  9. Geoffrey C. Bowker and Susan Leigh Star, Sorting Things Out: Classification and Its Consequences (MIT Press, 1999): on ‘torque’, and on the way every category valorizes one perspective and silences another.↩︎
  10. On the ‘displaced person’ of 1945 as a produced political category, Sebastian Huhn and Christoph Rass, ‘Displaced person(s): the production of a powerful political category’ (Ethnic and Racial Studies 48/4, 2025, 718–739). On ‘migrantization’, Katharine Charsley and Nicole Hoellerer, ‘Migrantisation: a key concept’ (Comparative Migration Studies 13, 2025); and Janine Dahinden, ‘A plea for the “de-migranticization” of research on migration and integration’ (Ethnic and Racial Studies 39/13, 2016, 2207–2225).↩︎
  11. Ahmet Celikten, Maik Hoops, Christoph Rass and Lale Yildirim, ‘Situating “Gastarbeiterkinder”: Science, School, and the Production of Figures of Migration’, in Production of Migration, vol. II: Empirical Approaches, ed. Isabelle Löhr et al. (Cham: Springer, forthcoming). The contribution shows that the city administration’s 1971 municipal report had deliberately avoided the compound ‘Gastarbeiterkind’. Primary source: Reinhardt Köhler, Bericht zur Situation der Gastarbeiter in Osnabrück (Stadt Osnabrück / Der Oberstadtdirektor, 1971), Stadtarchiv Osnabrück.↩︎
  12. John R. Searle, The Construction of Social Reality (Free Press, 1995), on institutional / social facts.↩︎
  13. J. B. Harley, ‘Deconstructing the Map’ (Cartographica 26/2, 1989, 1–20); Johanna Drucker, ‘Humanities Approaches to Graphical Display’ (Digital Humanities Quarterly 5/1, 2011), on ‘capta’; Henk van Houtum and Rodrigo Bueno Lacy, ‘The Migration Map Trap. On the Invasion Arrows in the Cartography of Migration’ (Mobilities 15/2, 2020, 196–219).↩︎
  14. Corpus of c. 6,000 German-language texts on migration and the school system: roughly 70 million words in raw form, c. 45 million in the cleaned working corpus (the figure used here).↩︎
  15. José van Dijck, ‘Datafication, Dataism and Dataveillance: Big Data between Scientific Paradigm and Ideology’ (Surveillance & Society 12/2, 2014).↩︎
  16. Michel-Rolph Trouillot, Silencing the Past: Power and the Production of History (Beacon Press, 1995).↩︎
  17. Bowker and Star, Sorting Things Out: on the way every category valorizes one perspective and silences another.↩︎
  18. Lara Putnam, ‘The Transnational and the Text-Searchable: Digitized Sources and the Shadows They Cast’ (American Historical Review 121/2, 2016).↩︎
  19. Ida Marie S. Lassen, Jens Christian Bjerring and Kristoffer L. Nielbo, ‘Silencing in Data Science Practices’ (Big Data & Society 12/3, 2025), which carries the vocabulary of silence into data work; contrasted here with Trouillot and with Ann Laura Stoler, Along the Archival Grain: Epistemic Anxieties and Colonial Common Sense (Princeton UP, 2009).↩︎
  20. Mrinalini Luthra, Konstantin Todorov, Charles Jeurgens and Giovanni Colavizza, ‘Unsilencing Colonial Archives via Automated Entity Recognition’ (Journal of Documentation 80/5, 2024, 1080–1105).↩︎
  21. Jessica Marie Johnson, ‘Markup Bodies: Black [Life] Studies and Slavery [Death] Studies at the Digital Crossroads’ (Social Text 36/4, 2018, 57–79).↩︎
  22. Processing of the Osnabrück indices: external service provider, deep-neural-network recognition exploiting the geometry of the structured forms, with a historical register of residents and street names supplied as a lexicon. On the data-driven processing of these mass card indexes, including the Osnabrück Gestapo-Kartei: Sebastian Bondzio and Christoph Rass, ‘Data Driven History – methodische Überlegungen zur Osnabrücker Gestapo-Kartei als Quelle zur Erforschung datenbasierter Herrschaft’, Archiv-Nachrichten Niedersachsen 22 (2018): 124ff.; and Sebastian Bondzio and Christoph Rass, ‘Die Osnabrücker Gestapo-Datei’, Osnabrücker Mitteilungen 124 (Bielefeld: Verlag für Regionalgeschichte, 2019).↩︎
  23. The ‘CM/1’ files are Arolsen Archives collection 3.2.1 (IRO ‘Care and Maintenance’ Program, c. 1947–1952). Prof. Gernot A. Fink holds the chair of Pattern Recognition at TU Dortmund (document-image analysis and handwriting recognition).↩︎
  24. On source-bound knowledge graphs with provenance and graded certainty, the lecture pointed to a wiki-based knowledge-graph approach presented earlier at the summer school.↩︎
  25. Jo Guldi, The Dangerous Art of Text Mining: A Methodology for Digital History (Cambridge University Press, 2023); Torsten Hiltmann et al., ‘NER4all, or Context is All You Need: Using LLMs for Low-Effort, High-Performance NER on Historical Texts. A Humanities-Informed Approach’ (arXiv:2502.04351, 2025).↩︎
  26. Transfer project ‘Reflexive Migration Research in the Museum’, with DOMiD (Dokumentationszentrum und Museum über die Migration in Deutschland), Cologne. DOMiD’s planned museum, long under the working title ‘Haus der Einwanderungsgesellschaft’, is to open as ‘Museum Selma’ (Cologne, scheduled 2029).↩︎
  27. Adrian Favell, The Integration Nation. Immigration and Colonial Power in Liberal Democracies (Polity, 2022); cf. Philosophies of Integration (Macmillan, 1998). The formulation ‘every analytical vocabulary tends to acquire the governance functions it set out to expose’ is the author’s own sharpening of Favell’s critique, not a quotation.↩︎
  28. As their supply of web text is exhausted, firms training large language models have turned to printed books as a data source, in some cases buying and destructively scanning them at scale. On these acquisitions and the antiquarian-book trade they have drawn on, see, e.g., ‘Inside an AI start-up’s plan to scan and dispose of millions of books’, The Washington Post, 27 January 2026; and ‘Rare book dealers fear tech firms are destroying obscure editions to train AI models’, NL Times, 25 June 2026.↩︎
  29. On domain-trained recognition of difficult historical hands, see the Swedish National Archives’ (Riksarkivet) ‘Swedish Lion’ handwritten-text-recognition model for Swedish court records, c. 1600–1900, developed by its AI lab (AIRA) with Transkribus and partner archives. On large language models reaching state-of-the-art transcription of historical manuscripts, see ‘Unlocking the Archives: Using Large Language Models to Transcribe Handwritten Historical Documents’ (arXiv:2411.03340, 2024). The claim here is that such models read difficult hands at scale, not that they surpass expert human readers in general.↩︎