Toward Universal Translatability: On the Statistical Ontology of Generative AI
American Comparative Literature Association 2026, Montreal. Seminar: Alphabet & AI
Introduction
My broader research on ML-based AI operates on two levels: one is ML as a discourse, and the second is ML as a set of concrete programming practices. The two levels differ substantially. Today, I focus on ML as a discourse, one that, I will argue, carries significant and largely unexamined philosophical commitments.
The hypothesis I explore in this talk is that translation, as a problem, marks both a point of origin of current generative AI and its normative horizon: the ambition for generative models to become so-called “world models.”
A model is an abstract, structured representation of an object. What characterizes such a representation and differentiates it from, say, a painting or a photograph is that a model is operative in the sense that it enables us not only to understand something about the object but also to very concretely act with it or upon it. A city map is a model in the sense that it enables us to orient ourselves in a city both cognitively and physically.
For generative models to become world models able to predict not solely the most likely sequence of tokens given an input but also the most likely next state of the world, they need to satisfy several conditions. And this is my own preliminary mapping of current discourse.
First, such models would require internal consistency — a stable and unified representation of their object. This is what current LLMs conspicuously lack, despite their capacity to produce fluent and contextually appropriate outputs (Vafa et al. 2024; Akyürek et al. 2024; Wolfram and Schein 2025).
Second, models would require consistency across data modalities — that is, structural alignment between the representation spaces of language and image. Some researchers have termed this the “Platonic Representation Hypothesis” (Huh et al. 2024; Jha et al. 2025): “Images (X) and text (Y) are projections of a common underlying reality (Z)” (Huh et al. 2024, 1).[1]
Third, it has been shown that, while LLMs infer certain physical properties solely from language (Abdou et al. 2021), they do not achieve a consistent understanding of physical reality. Models would need to capture and learn physical properties, which would allow an agent such as a robot to navigate its environment (Li 2025).
My hypothesis is that these conditions converge on a single requirement: that models become universal translation machines, that is, systems capable of establishing correspondences and mutual translatability across heterogeneous representation spaces and modalities: linguistic, visual, and physical. And the way that ML as a discourse frames this translatability reflects a metaphysical commitment:
Instead of understanding such a system of correspondences as the product of human-machine interaction and labor, its possibility is premised on the notion of an underlying statistical structure of reality—one that may be imperceptible to humans, yet that is accurately captured by models when they are exposed to enough data. I propose to trace this metaphysical commitment—which I tentatively term “statistical ontology” and which I infer from my reading of current ML research—back to Warren Weaver's 1949 Memorandum on Machine Translation and his envisioned solution to the problem of translation.
Short Genealogy of Machine Translation
In this short text, Weaver hopes for the existence of a “real but as yet undiscovered universal language” (23) that would be like an “open basement,” common to every linguistic edifice; a foundation allowing for an unencumbered communication between languages.
Therefore, what stands as the very beginning of machine translation—as a project that would take decades to yield even the smallest results—is the notion that languages can be understood in terms of common structures or “invariant properties” (16), decipherable like a code, with semantics emerging from mathematical-statistical structures. Weaver writes: “. . . it is very tempting to say that a book written in Chinese is simply a book written in English which was coded into the ‘Chinese code’“ (22).[2]
Drawing an analogy between translation and code deciphering, Weaver recounts the case of a mathematician who succeeds in cracking a language as though it were a code, without any semantic understanding of the message: “The most important point, at least for the present purposes, is that the decoding was done by someone who did not know Turkish, and did not know that the message was in Turkish” (Weaver 1955, 16).
Through this cryptographic framing, language is reconceived primarily as a code whose syntax is amenable to the basic operations of “scanning,” “writing,” “combining,” and “substituting”—which is precisely what computers, defined as Universal Turing Machines, are designed to do. By reconceiving translation as a purely syntactic operation, Weaver was the first to propose that machine translation is a “formally solvable” problem (22).
Today’s large language models seem to accomplish the unthinkable in Weaver’s time: the “unimpeded” access to source texts of any kind, including sources in minority languages with very limited training data. With the advent of neural networks, models are trained to capture relationships between words and encode dependencies beyond the immediate vicinity of the processed word.
This approach was initially based on “sequence modeling techniques” the purpose of which was to solve the so-called “structured output problem,” a problem that emerges when the structure of the output must be “related to the structure of the input” (Cho et al. 2015, p. 1875). In translation problems, both the input and the output are syntactically complex, with each word or token contextually dependent on the others; the question, then, is how to map one structure onto another. The method consists in mapping the probability distribution of the subcomponents of the output to the one of the subcomponent of the input.
From that point on, translation is redefined as the operation of mapping and aligning probability distributions. Linguistic translation becomes just one of the many tasks that can be addressed through sequence modeling. Automated image caption generation, for instance, relies on the same method. It consists in aligning the spatiality of the image or the spatio-temporality of the video with the linearity of the caption in order “to accurately describe the spatial relationships between elements of the scene represented in the image” (Cho et al. 2015, p. 1875). Through sequence modeling, translation becomes a matter of stabilizing relations of correspondence.
In 2017, in the field of so-called image translation, models were developed that learn to render a photograph in the style of a given painter (Zhu et al. 2017). As one research group puts it, “in analogy to automatic language translation, we define automatic image-to-image translation as the problem of translating one possible representation of a scene into another, given sufficient training data” (Isola et al. 2017). Here again, translation designates the transposition of an underlying spatial structure from one visual style to another—the structure or representational content of the image remaining invariant across the transformation.
The translation capabilities of recent LLMs are a byproduct of their generalization abilities, that is, their capacity to process unseen data. Researchers have shown that the structural representations a LLM acquires from one language can transfer to other languages without requiring equivalent quantities of training data for each (Pires, Schlinger, and Garrette 2019; Libovický, Rosa, and Fraser 2019). However, the authors hypothesize that language-invariant elements such as numbers and URLs function as anchors that contribute to the implicit alignment between token sequences across languages. This suggests that the cross-lingual correspondences observed in these models may be a product of shared features in the training data rather than emergent structural properties.
More recently, researchers have reported evidence that models develop concept representations that are independent of specific languages (Dumas et al. 2025). One of the papers is tellingly titled: “Separating tong from thought.” This echoes Weaver's hope that all languages would contain “basic common characteristics,” analogous to what he calls “tree-ness”: the abstract property, or in ML language the set of features, by which trees of different species are recognized as belonging to the same category. What is implicitly assumed here is that models would acquire as an emergent property, through training alone, something like a schema in the Kantian sense — that is, a mediating structure capable of bridging heterogeneous representations. A full critical analysis of this claim lies beyond the scope of this talk.
The Horizon: World Models as Universal Translation Machines
Translation becomes, in this framework, the stabilization of correspondences and equivalences between heterogeneous domains through the alignment between probability distributions—with the aspiration of moving between these domains seamlessly and without loss.
This understanding of translation differs from those developed in the humanities. Since the early 2000s, Translation Studies has been attentive not only to the power dynamics between languages (the colonial vs. minority language binary, for instance) but also to the translator’s positionality as a subject who speaks from a specific cultural, historical, and political location. The translator, as a political subject who guides readers across the liminal threshold between languages, cultural contexts, and temporalities, is a potential disruptor of hegemonic orders. That potential for resistance disappears with current machine translation and its implicit neo-colonial assumptions: that the opacities of languages are problems to be technically resolved, rather than generative markers of cultural difference.
Within the statistical ontology of ML-based AI, nothing in principle escapes translatability, that is, reduction to a universal system of computable equivalences. In other words, there is no exteriority that world models can’t enclose.
This statistical ontology reintroduces and naturalizes a realist metaphysics—one in which the cultural, historical, and economic factors that shape linguistic relations, technological development, and scientific research are rendered invisible, while world models are naturalized as the capture and reflection of an underlying structure of the real. What I have termed elsewhere the “normative rationality” of generative AI—the fact that, unlike previous symbolic AI, current AI is trained to automates human values, judgments, and interpretations—is thereby foreclosed.
What is also emerging through this statistical ontology is a renewed ontologizing of the relation between language, image, and reality. It is the assumption that language and image are mappings of the real, and that such mappings leave the real itself unchanged. This is a pre-Kantian, pre-critical metaphysics in which the noumenal world can be accessed and known as such. It renders invisible the vast infrastructure of human labor that made this series of correspondences learnable in the first place — the parallel corpora, the bilingual dictionaries, the multilingual translations of novels and poems, the ImageNet dataset, the annotation labor that made possible everything these models are claimed to achieve. Statistical ontology, in short, forecloses what is in fact a massive human endeavor.
Endnotes
[1] “Our conjecture is as follows: neural networks trained with the same objective and modality, but with different data and model architectures, converge to a universal latent space such that a translation between their respective representations can learned without any pairwise correspondence” (Jha et al. 2025, 3).
[2] It would be necessary to review the connection between Weaver’s conception of translation and Searl’s Chinese Room.
References
Abdou, Mostafa, Artur Kulmizev, Daniel Hershcovich, Stella Frank, Ellie Pavlick, and Anders Søgaard. 2021. “Can Language Models Encode Perceptual Structure Without Grounding? A Case Study in Color.” arXiv:2109.06129. Preprint, arXiv, September 14. https://doi.org/10.48550/arXiv.2109.06129.
Akyürek, Afra Feyza, Ekin Akyürek, Leshem Choshen, Derry Wijaya, and Jacob Andreas. 2024. “Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability.” arXiv:2401.08574. Preprint, arXiv, June 26. https://doi.org/10.48550/arXiv.2401.08574.
Dumas, Clément, Chris Wendler, Veniamin Veselovsky, Giovanni Monea, and Robert West. 2025. “Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in Transformers.” Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics 1 (July): 31822–41. https://doi.org/10.18653/v1/2025.acl-long.1536.
Huh, Minyoung, Brian Cheung, Tongzhou Wang, and Phillip Isola. 2024. “The Platonic Representation Hypothesis.” Proceedings of the 41 St International Conference on Machine Learning 235 (July): 1–26. https://proceedings.mlr.press/v235/huh24a.html.
Jha, Rishi, Collin Zhang, Vitaly Shmatikov, and John X. Morris. 2025. “Harnessing the Universal Geometry of Embeddings.” arXiv:2505.12540. Preprint, arXiv, May 20. https://doi.org/10.48550/arXiv.2505.12540.
Lerchner, Alexander. 2026. “The Abstraction Fallacy: Why AI Can Simulate But Not Instantiate Consciousness.” Preprint, PhilPaper, March 10. https://deepmind.google/research/publications/231971/.
Li, Fei-Fei. 2025. “From Words to Worlds: Spatial Intelligence Is AI’s Next Frontier.” Substack newsletter. Dr. Fei-Fei Li, November 10. https://drfeifei.substack.com/p/from-words-to-worlds-spatial-intelligence.
Vafa, Keyon, Justin Y. Chen, Ashesh Rambachan, Jon Kleinberg, and Sendhil Mullainathan. 2024. “Evaluating the World Model Implicit in a Generative Model.” arXiv:2406.03689. Preprint, arXiv, November 10. https://doi.org/10.48550/arXiv.2406.03689.
Wolfram, Christopher, and Aaron Schein. 2025. “World Models and Consistent Mistakes in LLMs.” Paper presented at ICML 2025 Workshop on Assessing World Models. ICML.