Contents
Proto-Indo-European is the reconstructed ancestral language of the Indo-European language family, spoken during the Late Neolithic and Early Bronze Age. Because no written documents survive, linguists reconstructed its vocabulary, phonology, and grammar through comparative and internal reconstruction. It was a fusional language featuring pitch accent, vowel ablaut, eight or nine nominal cases, and a complex verbal system structured around grammatical aspect. Most scholars associate its speakers with the pastoralist Yamnaya culture of the Pontic-Caspian steppe, whose migrations spread regional dialects across Eurasia, giving rise to the historical Indo-European daughter languages.
History of scholarship
Observations of recurring similarities between European and Asian languages date to the early modern period. In the sixteenth century, European travelers and missionaries visiting the Indian subcontinent noted striking correspondences between Indian tongues and European languages. In 1647 and 1653, the Dutch scholar Marcus Zuerius van Boxhorn formulated a hypothesis that Germanic, Romance, Greek, Baltic, Slavic, Celtic, and Iranian descended from a shared primitive ancestor that he called Scythian. In 1767, Gaston-Laurent Coeurdoux, a French Jesuit missionary living in India, sent a memoir to the Academie des Inscriptions et Belles-Lettres in Paris demonstrating regular correspondences between Sanskrit and Latin and Greek. In 1786, the Anglo-Welsh philologist William Jones delivered his famous third anniversary discourse to the Asiatic Society in Calcutta, noting that Sanskrit, Greek, and Latin displayed structural affinities in verb roots and grammatical forms too strong to be accidental, and suggesting that Gothic, Celtic, and Persian sprang from the same source. Jones, however, erroneously grouped Egyptian, Japanese, and Chinese into the family while omitting Hindi.
The systematic methodology of historical-comparative linguistics took shape in the early nineteenth century. In 1816, Franz Bopp published his comparative study of the conjugation systems of Sanskrit, Greek, Latin, Persian, and Germanic, demonstrating their genetic relationship and establishing comparative grammar as an academic discipline. Independently, Danish linguist Rasmus Rask completed an essay in 1814, published in 1818, demonstrating the relation of Germanic to Baltic, Slavic, Greek, and Latin. In 1822, Jacob Grimm formulated Grimm's law in the second edition of his Deutsche Grammatik, establishing that consonant shifts in Germanic occurred systematically and regularly. August Friedrich Pott subsequently published comprehensive tables of phonetic correspondences across major Indo-European branches between 1833 and 1836, providing an etymological foundation for the discipline.
The first comprehensive reconstruction of Proto-Indo-European grammar was published by August Schleicher in 1861 in his Compendium der vergleichenden Grammatik der indogermanischen Sprachen. Schleicher introduced the genealogical tree model (Stammbaumtheorie) and used asterisks to mark reconstructed, unattested forms. To illustrate the state of reconstruction, he composed a fable in the reconstructed proto-language, titled The Sheep and the Horses. Schleicher relied heavily on Sanskrit, which he considered closest to the ancestral language, and reconstructed a system with only three short vowels (a, i, u) and their long counterparts. In 1868, August Fick published the first comparative etymological dictionary of the Indo-European languages, assembling a vast corpus of reconstructed roots.
During the 1870s, the Neogrammarian school emerged at the University of Leipzig, led by August Leskien, Karl Brugmann, Hermann Osthoff, Berthold Delbruck, and Hermann Paul. The Neogrammarians established the principle that sound laws operate mechanically without exceptions within the same dialect and environment, and that apparent irregularities resulted from analogy or the interference of competing phonetic environments. This principle was reinforced by Karl Verner's 1876 discovery of Verner's law, which explained apparent exceptions to Grimm's law through the position of the original Proto-Indo-European pitch accent. In the late 1870s, the discovery of the law of palatals in Indo-Iranian proved that Proto-Indo-European originally distinguished e and o, demonstrating that Sanskrit had merged original e, o, and a into a single vowel, making its vowel system innovative rather than archaic.
A major conceptual advance occurred in 1878 when the young Swiss linguist Ferdinand de Saussure published his treatise on the primitive vowel system of Indo-European. Using internal reconstruction, Saussure posited two theoretical entities that he termed sonantic coefficients, which did not survive in documented languages but had conditioned vocalic ablaut and lengthened adjacent vowels upon disappearing. In 1880, Hermann Moller added a third coefficient and proposed that they were guttural or laryngeal consonants. Although the Neogrammarians initially rejected Saussure's proposal, it was vindicated in 1927 when Polish linguist Jerzy Kurylowicz demonstrated that the consonant h in newly deciphered Hittite occurred in positions predicted by Saussure's hypothetical coefficients. This confirmation led to the development of the laryngeal theory.
In the twentieth century, comparative linguistics incorporated new discoveries, notably the Anatolian and Tocharian branches, which forced revisions of long-held assumptions regarding dialect grouping and archaic features. Antoine Meillet emphasized that the proto-language was not a uniform, indivisible entity but a complex continuum containing dialectal variation that could only be approximated as a system of regular correspondences. Jerzy Kurylowicz's 1956 monograph on ablaut and Emile Benveniste's work on nominal morphology advanced structural understanding. Julius Pokorny's 1959 Indogermanisches etymologisches Worterbuch summarized lexical knowledge accumulated up to that time, while Helmut Rix and his collaborators systematized verb roots in the Lexikon der indogermanischen Verben. In the late twentieth century, alternative paradigms emerged, such as the glottalic theory proposed by Thomas Gamkrelidze, Vyacheslav Ivanov, and Paul Hopper, and typological studies of syntax by Winfred Lehmann.
Classification and daughter languages
The breakup of Proto-Indo-European produced ten recognized primary branches, alongside several fragmentarily attested languages. The oldest attested branch is Anatolian, preserved on cuneiform tablets from the eighteenth to twelfth centuries BCE in central and western Anatolia. Anatolian includes Hittite, Luwian, Palaic, Lycian, Lydian, and Carian. Hittite preserves direct reflexes of two laryngeals, retains a two-gender system of common and neuter, and lacks several morphological categories found elsewhere in the family, leading Edgar Sturtevant and later scholars to propose the Indo-Hittite hypothesis, which views Anatolian as having separated prior to the divergence of all other branches.
The Tocharian branch consists of Tocharian A (Turfanian) and Tocharian B (Kuchean), documented in Buddhist religious and administrative manuscripts from the sixth to eighth centuries CE in the Tarim Basin of modern Xinjiang, China. Despite their eastern geographic location, both Tocharian languages display centum phonetic developments, preserving labiovelars while merging palatovelars with plain velars, a fact that dismantled the nineteenth-century belief that centum and satem languages represented a simple western versus eastern geographic division.
The Italic branch encompasses Latin, Faliscan, Oscan, Umbrian, and several minor dialects of ancient Italy. Latin expanded across Europe and the Mediterranean Basin, eventually diversifying into the Romance languages, including Portuguese, Spanish, Catalan, Occitan, French, Italian, Romansh, and Romanian. The Celtic branch once stretched across continental Europe from the Iberian Peninsula to Galatia in Anatolia, represented by Gaulish, Celtiberian, and Lepontic, but survived only in the Insular Celtic languages of the British Isles and Brittany, divided into Goidelic (Irish, Scottish Gaelic, Manx) and Brittonic (Welsh, Breton, Cornish). Proposals grouping Italic and Celtic into an Italo-Celtic clade remain widely discussed.
The Germanic branch divided into West Germanic (English, German, Dutch, Frisian, Yiddish, Afrikaans), North Germanic (Old Norse, Icelandic, Faroese, Norwegian, Danish, Swedish), and East Germanic (Gothic, Vandalic, Burgundian, Crimean Gothic, all now extinct). Germanic is characterized by the First Germanic Consonant Shift described by Grimm's and Verner's laws, the fixation of stress accent on initial root syllables, and the reduction of verbal categories into strong and weak conjugations. The Balto-Slavic branch branched into the Baltic group (Lithuanian, Latvian, Old Prussian) and the Slavic group, which further subdivided into East Slavic (Russian, Ukrainian, Belarusian, Rusyn), West Slavic (Polish, Czech, Slovak, Sorbian, Kashubian), and South Slavic (Bulgarian, Macedonian, Serbo-Croatian, Slovene). Baltic and Slavic share extensive innovations in accentuation, morphology, and lexicon, supporting their descent from a common Proto-Balto-Slavic node.
The Indo-Iranian branch comprises three groups: Indo-Aryan, Iranian, and Nuristani. Indo-Aryan is attested from the mid-second millennium BCE in the sacred hymns of the Rigveda, developing through classical Sanskrit and Middle Indo-Aryan Prakrits into modern languages such as Hindi-Urdu, Bengali, Punjabi, Marathi, and Gujarati. Iranian is documented in Old Avestan from the early first millennium BCE and Old Persian royal inscriptions, ancestral to Persian, Pashto, Kurdish, Balochi, and Ossetian. Nuristani languages are spoken in the remote valleys of eastern Afghanistan and northwestern Pakistan. Indo-Iranian shares profound lexical, mythological, and morphological innovations with Greek and Armenian, supporting proposed Graeco-Aryan and Graeco-Armenian connections.
The Hellenic branch is represented by Greek, which boasts the longest continuous written record among living Indo-European languages, beginning with Mycenaean Greek recorded in Linear B tablets from the fourteenth to thirteenth centuries BCE and continuing through ancient dialects to Modern Greek and Tsakonian. Armenian represents an independent branch attested from the fifth century CE in Classical Armenian (Grabar). Albanian forms an isolated branch with written records dating from the fifteenth century CE, though its vocabulary preserves deep archaic elements alongside heavy Latin, Greek, Slavic, and Turkish loans. Several poorly attested ancient languages of southeastern Europe and the Mediterranean, known collectively as Paleo-Balkan languages, include Phrygian, Thracian, Dacian, Illyrian, Messapic, Venetic, and Lusitanian, among which Phrygian shows particularly close affinities with Greek.
Phonology
Proto-Indo-European phonology has been reconstructed with significant detail. The standard consonant inventory consists of stops, nasals, liquids, semivowels, a sibilant fricative, and three laryngeals. The sonorants could function either as consonantal syllable margins or as syllabic vocalic nuclei when positioned between other consonants or word-initially before consonants. The language possessed a distinctive vowel ablaut system and a free, mobile pitch accent.
Stop consonants and the glottalic theory
In the classical reconstruction associated with nineteenth-century Neogrammarians such as Karl Brugmann, the language was assigned four series of stops across five points of articulation: voiceless (*p, *t, *ḱ, *k, *kʷ), voiceless aspirates (*pʰ, *tʰ, *ḱʰ, *kʰ, *kʷʰ), voiced (*b, *d, *ǵ, *g, *gʷ), and voiced aspirates (*bʰ, *dʰ, *ǵʰ, *gʰ, *gʷʰ). In 1891, Ferdinand de Saussure demonstrated that voiceless aspirates were secondary formations arising from clusters of a plain voiceless stop and a following laryngeal. Consequently, most twentieth-century linguists adopted a three-series system without voiceless aspirates: voiceless, voiced, and voiced aspirates.
The traditional three-series system faced two prominent typological criticisms. Roman Jakobson observed that languages possessing voiced aspirated stops almost universally possess voiceless aspirated stops as well, making the reconstructed system typologically abnormal, although Kelabit in Borneo has been cited as a rare modern counterexample. Furthermore, Holger Pedersen noted that the voiced labial stop *b is extraordinarily rare in reconstructed roots, whereas across natural languages with missing stops, it is typically voiceless *p that is absent rather than voiced *b. To address these problems, Nikolai Andreev suggested in 1957 that Proto-Indo-European stops differed in strength rather than voicing, positing fortis, lenis, and aspirated fortis series analogous to Korean.
In 1972 and 1973, Thomas Gamkrelidze, Vyacheslav Ivanov, and Paul Hopper independently proposed the glottalic theory. Under this framework, the traditional voiced series (*b, *d, *ǵ, *g, *gʷ) was reinterpreted as voiceless glottalized stops or ejectives (*p', *t', *ḱ', *k', *kʷ'), while the traditional voiceless and voiced aspirated series were reinterpreted as voiceless and voiced stops with non-distinctive aspiration. This model provided an explanation for the rarity of *b, since labial ejectives (*p') are typologically rare in ejective inventories. It also reformulated Grimm's law, Grassmann's law, and Bartholomae's law as retentions rather than complex sound shifts. Critics of the glottalic theory argue that the spontaneous voicing of ejective stops across almost all daughter branches is typologically unprecedented, especially in initial word positions, and that correspondences with Proto-Kartvelian ejectives do not match the expected sound classes.
Dorsal consonants and the centum-satem split
The traditional reconstruction posits three series of dorsal stops: palatovelars (*ḱ, *ǵ, *ǵʰ), plain velars (*k, *g, *gʰ), and labiovelars (*kʷ, *gʷ, *gʷʰ). Across the daughter languages, these three series merged into two in different configurations, forming the basis of the centum-satem division. In the centum languages (Italic, Celtic, Germanic, Hellenic, Anatolian, Tocharian), palatovelars merged with plain velars into plain velars, while labiovelars remained distinct. In the satem languages (Indo-Iranian, Balto-Slavic, Armenian, Albanian), labiovelars lost their labialization and merged with plain velars, while palatovelars developed into affricates or sibilant fricatives, as seen in the word for hundred: Latin centum versus Avestan satəm, reflecting Proto-Indo-European *ḱm̥tóm.
Because no single daughter language unequivocally preserves all three series independently across its entire vocabulary, several linguists challenged the three-dorsal model. Hermann Hirt, Antoine Meillet, and Aleksey Savchenko argued that the centum two-series system (velars and labiovelars) was original and that palatovelars arose secondary via fronting. Jerzy Kurylowicz suggested that the satem system was primitive. Stefan Mladenov and Jan Safarewicz proposed that Proto-Indo-European had only a single velar series that split conditionally. However, evidence from Luwian, Armenian, and Albanian indicates that certain environments distinguish reflexes of all three dorsal series, supporting the three-series reconstruction as representing the late ancestral language.
Fricatives and laryngeals
Proto-Indo-European possessed a single unquestioned sibilant fricative, *s, which exhibited an automatic voiced allophone, [z], when positioned immediately before a voiced consonant, as in *ni-sd-ós ('nest', literally 'sitting down', derived from the root *sed-). Karl Brugmann hypothesized four interdental fricatives (*þ, *þʰ, *ð, *ðʰ) to explain correspondences between Greek dental or labial stops and Sanskrit sibilants, but subsequent research demonstrated that these reflexes represented original clusters of a dental stop followed by a dorsal stop (*TK). Gamkrelidze and Ivanov suggested adding palatalized *ś and labialized *śʷ, though this proposal has not achieved broad acceptance.
The laryngeal consonants, conventionally transcribed as *h₁, *h₂, and *h₃, form a central pillar of modern Indo-European phonology. Although estimates of the number of laryngeals have ranged from one to ten, the three-laryngeal consensus is prevailing. Their precise phonetic values remain debated: *h₁ was likely a neutral glottal stop [ʔ] or glottal fricative [h]; *h₂ was a voiceless uvular or pharyngeal fricative ([χ] or [ħ]); and *h₃ was a voiced labialized velar, uvular, or pharyngeal fricative ([ɣʷ], [ʁʷ], or [ʕʷ]). When adjacent to vowels, laryngeals conditioned quality changes: *h₁ produced no color change (*h₁e > *e), *h₂ colored *e to *a (*h₂e > *a), and *h₃ colored *e to *o (*h₃e > *o). Following a short vowel, an adjacent laryngeal was lost with compensatory lengthening: *eh₁ > *ē, *eh₂ > *ā, *eh₃ > *ō. Between consonants, laryngeals vocalized as vowels in the daughter languages, yielding Greek e, a, o; Indo-Iranian i; and a in most other branches.
Vowels and sonorants
The reconstructed vowel inventory comprises short *e, *o, and arguably *a, alongside long *ē, *ō, and long *ā. The high vowels *i and *u functioned as the syllabic vocalic allophones of the semivowels *y and *w, occurring in nuclear syllable position between consonants. Most instances of long vowels *ī and *ū, as well as *ā, arose through compensatory lengthening following the loss of adjacent laryngeals (*iH > *ī, *uH > *ū, *eh₂ > *ā). Diphthongs combined short or long e, o, and a with the semivowels: *ey, *oy, *ay, *ew, *ow, *aw, and the rarer lengthened-grade diphthongs *ēy, *ōy, *ēw, and *ōw.
The sonorants *m, *n, *r, and *l exhibited dual functionality. Preceding or following vowels, they operated as consonants; positioned between consonants or word-initially before a consonant, they assumed syllabic status, functioning as vocalic nuclei (*m̥, *n̥, *r̥, *l̥). Sanskrit alone among historical daughter languages preserved syllabic r̥ as a distinctive vowel. In other branches, syllabic sonorants developed epenthetic vowels, producing reflexes such as Latin em, en, or, ol; Greek a, ar, al; Germanic um, un, ur, ul; and Slavic im, in, ir, il. When a root began with a cluster of two stops followed by a resonant, an epenthetic reduced vowel known as schwa secundum (*ₑ or *ₔ) developed between the stops to facilitate articulation, as seen in *kʷₑtwor- ('four').
Accent and prosody
Proto-Indo-European possessed a free, mobile pitch accent that could fall on any syllable of a word and shift positions across the declensional or conjugational paradigm of a single lexeme. Accented syllables were pronounced with higher pitch. The accent is best preserved in Vedic Sanskrit and Ancient Greek, with significant indirect reflexes preserved in Balto-Slavic tonal accents and in Germanic through Verner's law, where voiceless fricatives became voiced when following an unaccented syllable. Accent placement was lexically distinctive: in Greek, phóros ('tribute') with root accent contrasts with phorós ('bearing') with suffix accent, and tróchos ('course') contrasts with trochós ('wheel').
Scholars categorize athematic nominals and verbs into accent-ablaut classes based on the distribution of stress and vowel grades between strong cases (nominative, accusative, vocative) and weak cases (genitive, dative, ablative, locative, instrumental). Holger Pedersen established the distinction between proterokinetic stems, which accented the root in strong cases and the suffix in weak cases, and hysterokinetic stems, which accented the suffix in strong cases and the ending in weak cases. Later linguists, including James Mallory, Douglas Adams, and Michael Meier-Brugger, identified acrostatic stems, which maintain fixed stress on the root throughout the paradigm, and amphikinetic stems, which stress the root in strong cases and the ending in weak cases. Thematic stems generally exhibited fixed stress, while finite verbs were typically enclitic (unaccented) in main clauses but accented in subordinate clauses.
Morphophonology and ablaut
The morphophonology of Proto-Indo-European was organized around the root, typically of the shape CVC, with consonants at both margins and an inherent vowel *e. Roots could be expanded to CCVC, CVCC, or CCVCC, and optionally preceded by mobile *s- (s-mobile). Roots were subject to strict phonotactic constraints: two voiced stops could not occur in the same root (*DeD was prohibited); a voiceless stop could not co-occur with a voiced aspirated stop (*TeDʰ and *DʰeT were prohibited); and two identical stops were barred (*tet-, *kek-).
Indo-European ablaut (apophony) was a regular system of vowel variation affecting roots, suffixes, and endings. Ablaut operated in two dimensions: qualitative variation between *e and *o, and quantitative variation between full grade (*e, *o), lengthened grade (*ē, *ō), and zero grade (complete absence of the vowel, Ø). In roots containing semivowels or sonorants, the zero grade caused the resonant to vocalize (*ey > *i, *ew > *u, *er > *r̥). While every morpheme theoretically possessed five ablaut grades, in practice particular morphological categories selected specific grades: present stems typically took the full e-grade, perfect singular stems took the o-grade, and aorists or weak nominal cases took the zero grade.
Nominal morphology
Nominals (nouns and adjectives) inflected for case, number, and gender. The late proto-language distinguished three grammatical genders: masculine, feminine, and neuter. Following Antoine Meillet, historical linguists widely accept that this tripartite gender system developed from an earlier two-class system distinguishing animate (common) and inanimate (neuter) nouns, as preserved in the Anatolian languages. The feminine gender arose as a late innovation outside Anatolian, emerging through the specialization of an inanimate collective suffix in *-h₂, which yielded feminine nouns in *-eh₂ (becoming -ā in daughter branches) and adjectives in *-ih₂- or *-yeh₂-.
The language distinguished three grammatical numbers: singular, dual, and plural. The dual referred specifically to a pair of entities, such as eyes, hands, or two individuals acting together, but was gradually lost across most daughter branches. Inanimate neuter nouns did not distinguish nominative, accusative, and vocative forms, using an unmarked zero-ending in the singular and the collective suffix *-h₂ in the plural. Plural neuter collective subjects routinely took singular verbs, a syntactic rule preserved intact in Ancient Greek (panta rhei, 'all things flows') and archaic Latin (pecunia non olet).
Proto-Indo-European possessed eight canonical cases: nominative (subject), vocative (direct address), accusative (direct object and goal of motion), genitive (possession and partitive relation), ablative (source and movement away), dative (indirect object and recipient), instrumental (means, tool, or accompaniment), and locative (place where and time when). Old Hittite preserves traces of a ninth case, the allative or directive, marking direction toward an object, with a tentative ending in *-ō or *-a. Ancient Indo-Iranian languages retained all eight cases, while other branches progressively syncretized them: Old Church Slavonic and Baltic retained seven (merging the ablative with the genitive), Latin six, Ancient Greek five, and Gothic four.
Nominals were divided into thematic stems, which inserted a thematic vowel (*-o- alternating with *-e-) between the stem and the case ending, and athematic stems, which attached endings directly to the root or derivational suffix. Athematic endings were characterized by ablaut and mobile accent, whereas thematic nouns showed fixed accent and contracted endings. Consonantal athematic stems included root nouns (*pōds, genitive *pedés, 'foot'), r-stems (*ph₂tḗr, 'father'), n-stems (*h₁néh₃mn̥, 'name'), s-stems (*nébʰos, 'cloud'), and heteroclitic neuters that alternated between *-r- in the nominative-accusative and *-n- in oblique cases, such as *yēkʷr̥ ('liver'), genitive *yekʷnós, and *wódr̥ ('water'), genitive *wednós.
Case endings reconstructed by Andrew Sihler, Donald Ringe, Miguel Villanueva Svensson, and Benjamin Fortson show broad agreement with minor variations. In the singular, athematic nominative was *-s (or zero in resonant stems), accusative *-m̥, genitive *-os or *-es, ablative *-s (identical to genitive), dative *-ey, locative *-i or endingless, instrumental *-h₁ or *-eh₁, and vocative endingless. In thematic singulars, nominative was *-os, accusative *-om, genitive *-osyo, ablative *-ōd, dative *-ōy, locative *-oy or *-ey, instrumental *-oh₁ (contracting to *-ō), and vocative *-e. In the dual, nominative-accusative-vocative was *-h₁e for athematic and *-oh₁ for thematic stems.
In plural oblique cases, daughter languages diverge into two distinct areal groupings known as the m-languages and the bh-languages. Baltic, Slavic, and Germanic exhibit dative-ablative plural in *-mos and instrumental plural in *-mis, whereas Indo-Iranian, Greek, Italic, and Celtic display dative-ablative in *-bʰos and instrumental in *-bʰis. Hittite alone preserves an ending in *-os without either consonant, suggesting that *-bʰ- and *-m- were originally independent postpositional particles that grammaticalized into case affixes in separate regional dialects of late Proto-Indo-European.
Adjectives and the Caland system
Adjectives inflected identically to nouns across case, number, and gender, agreeing with the head noun they modified. Masculine and neuter forms typically followed the thematic *-o- declension, while feminine forms took the *-eh₂- declension. Athematic adjectives in *-u- or *-nt- formed their feminines using the suffix *-ih₂- or *-yeh₂-. In older stages prior to the evolution of the feminine gender, adjectives distinguished only animate and neuter forms (*sh₂eldus 'sweet', neuter *sh₂eldu).
A notable subset of roots formed adjectives governed by the Caland system, formulated by Dutch indologist Willem Caland. When roots of this class formed adjectives, they used zero-grade roots with suffixes such as thematic *-ro- (*h₁rudʰ-ró- 'red', *h₂rǵ-ró- 'bright'), athematic *-u-, or *-nt-. When appearing as the first member of a compound, these adjectives substituted the stem suffix with *-i- (Greek argi-keraunos, 'with bright lightning'). In verbal derivation, Caland roots regularly formed stative verbs taking the suffix *-eh₁- (Latin rubēre 'to be red').
Adjective comparison was accomplished through derivational suffixes. The comparative was marked by amphikinetic *-yos- (nominative *-yōs, genitive *-is-és, zero-grade *-is-), surviving in Latin maior ('greater') and Lithuanian naujesnis ('newer'), or contrastive *-(t)ero-, denoting one of two alternatives (Greek póteros 'which of the two', Lithuanian katras). The superlative was marked by *-isto- (combining zero-grade *-is- with *-to-, surviving in English -est and Greek -istos) or *-m̥mo- (surviving in Latin optimus and Sanskrit -tama-).
Pronouns
Proto-Indo-European pronouns had unique paradigms that differed substantially from nouns. Personal pronouns existed only for the first and second persons; demonstrative pronouns functioned in place of third-person pronouns. Personal pronouns lacked grammatical gender and exhibited suppletion between nominative and oblique stems. In the first-person singular, the nominative stem *h₁eǵ- (or *h₁eǵoH) contrasted with the oblique stem *h₁me-, preserved in English I versus me. In the plural, the first-person nominative was *wei ('we'), contrasting with oblique *n̥s- ('us'). Second-person singular was *tuH ('thou'), oblique *twe-; second-person plural was *yuH ('ye'), oblique *us-.
Personal pronouns distinguished stressed tonic forms and unaccented enclitic forms in the accusative, genitive, and dative singular: first-person accusative *h₁mé versus enclitic *h₁me, genitive *h₁méne versus *h₁moi, dative *h₁méǵʰio versus *h₁moi; second-person accusative *twé versus *te, genitive *tewe versus *toi, dative *tébʰio versus *toi. Dual personal pronouns are tentatively reconstructed as first-person nominative *weh₁ and second-person *yuh₁.
Demonstrative pronouns operated on two distance tiers: proximal *is, *(h₁)id, *(h₁)ih₂ ('this') and distal *so, *tod, *seh₂ ('that'). In the distal demonstrative, the nominative singular masculine *so and feminine *seh₂ were uninflected forms lacking the final *-s, while all other cases used the stem *to- (neuter nominative-accusative *tod, accusative masculine *tom). Oblique masculine and neuter singular forms featured a characteristic formative *-sm- (*tosmōd ablative, *tosmey dative, *tosmi locative), while feminine obliques inserted *-sy- (*tosyās genitive, *tosyey dative). This pronominal declension also supplied relative and interrogative pronouns (*kʷis 'who?', *kʷid 'what?'). A reflexive pronoun lacking nominative and number distinction (*swe accusative, *sewe genitive, *sebʰi dative) denoted an object identical to the subject.
Numerals
Proto-Indo-European employed a decimal numeral system. Cardinal numerals from one to ten are securely reconstructed: *h₁óynos (or *sem-) 'one', *dwóh₁ 'two', *tréyes 'three', *kʷetwóres 'four', *pénkʷe 'five', *s(w)éḱs 'six', *septḿ̥ 'seven', *h₃eḱtṓw (or *oḱtṓw) 'eight', *h₁néwn̥ 'nine', and *déḱm̥(t) 'ten'. Only numerals one through four inflected for case and gender. Numerals five through ten were indeclinable.
The numeral *sem- was the archaic word for 'one, together', surviving in Ancient Greek heis (from *sems), neuter hen (*sem), and English simple, while *h₁óynos originally meant 'single, solitary'. 'Two' (*dwóh₁) was an inherent dual noun. 'Three' had masculine *tréyes, neuter *tríh₂, and feminine *tisres. 'Four' had masculine *kʷetwóres, neuter *kʷetwṓr, and feminine *kʷétesres. When used as prefixal compounding elements, they assumed zero-grade forms: *sm̥- ('single-'), *dwi- ('two-'), *tri- ('three-'), and *kʷ(e)tru- ('four-').
Decades from twenty to ninety were formed with a suffix reflecting *-(d)ḱomt-, derived from *déḱm̥ ('ten'): *wīḱm̥t- ('twenty', literally 'two tens'), *trīḱomt- ('thirty'), *kʷetwr̥̄ḱomt- ('forty'), *penkʷēḱomt- ('fifty'), *s(w)eḱsḱomt- ('sixty'), *septm̥̄ḱomt- ('seventy'), *h₃eḱtō(u)ḱomt- ('eighty'), and *h₁newn̥̄ḱomt- ('ninety'). The word for 'hundred', *ḱm̥tóm, is generally analyzed as an abbreviation of *dḱm̥t-dḱm̥tóm ('tenth ten' or 'ten tens'), though Winfred Lehmann suggested it originally designated an indefinite large quantity. 'Thousand' was represented by *ǵʰéslo- in southern branches (Greek khilioi, Sanskrit sahasram) and *tus-dḱm̥ti ('swollen hundred') in northern branches (Germanic thousand, Slavic tysyacha, Baltic tūkstantis). Ordinals were derived using suffixes *-mo- (*pr̥h₃-mó- 'first'), *-(t)ero- (*h₂én-tero- 'second, other'), and *-tó- (*tr̥-tó- 'third').
Verbal morphology
The Proto-Indo-European verb was structured around grammatical aspect rather than relative tense. Three aspectual stems were distinguished: the imperfective (present), depicting progressive, ongoing, or iterative action; the perfective (aorist), depicting action viewed as a completed whole; and the stative (perfect), denoting a state of being resulting from a prior action. Verbs conjugated for two voices: active and mediopassive (denoting actions in which the subject was personally affected or an action performed on oneself). Four moods were grammaticalized: indicative, imperative, subjunctive (expressing will, anticipation, or future eventuality), and optative (expressing wish or potentiality).
Verbal endings were divided into primary endings, used in the present indicative and subjunctive, and secondary endings, used in the past tenses (imperfect, aorist) and optative. In the active voice, athematic primary singular endings were first-person *-mi, second-person *-si, third-person *-ti, while thematic presents used first-person *-oh₂, second-person *-esi, third-person *-eti. Plural active primary endings were *-mos, *-te, *-nti. Secondary endings lacked the final *-i deictic particle: first-person *-m, second-person *-s, third-person *-t; plural *-me, *-te, *-nt. Dual active endings included primary first-person *-wos and secondary *-we.
The perfect possessed an independent set of active endings derived from an archaic stative conjugation: singular first-person *-h₂e, second-person *-th₂e, third-person *-e; plural first-person *-me, second-person *-e, third-person *-ēr. Perfect stems typically featured root reduplication with the vowel *e and took the o-grade in the singular indicative and zero-grade in the plural (*wóyd-h₂e 'I know', *wid-mé 'we know', from the root *weyd- 'to see'). In Hittite, the endings of the mi-conjugation correspond to traditional primary active endings, while the ḫi-conjugation corresponds to the Proto-Indo-European perfect endings, prompting Jay Jasanoff to propose that early Proto-Indo-European possessed two parallel present conjugations: a *mi-present and a *h₂e-present.
Present stems were formed through diverse morphological processes, including root presents (*h₁és-ti 'is'), reduplicated presents (*stí-steh₂-ti 'stands'), nasal-infix presents inserting *-né- in strong forms and *-n- in weak forms (*yu-né-g-ti 'joins', from *yewg-), suffixes in *-yé/ó- (*bʰoh₂-yé-ti 'causes to shine'), and inchoative-iterative suffixes in *-sḱé/ó- (*pr̥-sḱé-ti 'asks'). Aorist stems were formed as root aorists (*dʰéh₁-t 'placed'), thematic aorists (*wéyd-e-t 'saw'), or sigmatic aorists adding suffix *-s- (*wḗkʷ-s-t 'spoke', with lengthened grade in the active singular). Menne past forms in Greek, Indo-Iranian, and Armenian featured the augment *(h₁)e-, a proclitic temporal adverb that prefixed to past indicative verb forms.
Non-finite verbal formations included an array of participles: active present and aorist participles formed with *-nt- (*h₁s-ónt- 'being'), active perfect participles with *-wōs- / *-us- (*weyd-wōs- 'having known'), and mediopassive participles with *-mHno- or *-mh₁no- (*bʰér-o-mh₁no- 'being carried'). Verbal adjectives were derived with *-tó- and *-nó-, which later supplied passive past participles in Germanic, Slavic, and Latin. No universal infinitive existed in the parent language; instead, daughter languages independently grammaticalized various case forms of abstract verbal nouns ending in *-ti-, *-tu- (supine in *-tum), *-men-, or *-dʰye-.
Syntax and typology
Proto-Indo-European was a synthetic, fusional language that relied primarily on inflectional affixes rather than rigid word order to mark syntactic relations. Subject and verb agreed in person and number, while adjectives and demonstratives agreed with head nouns in case, number, and gender. Unmarked main clauses placed the verb in final position, conforming to a default subject-object-verb (SOV) order. Jacob Wackernagel initially reconstructed an SVO order in 1892 based on Vedic evidence, but Hans Henrich Hock and Winfred Lehmann demonstrated that SOV represents the ancestral pattern, as reflected in Old Indo-Aryan, Old Iranian, Old Latin, and Hittite.
Enclitic pronouns, conjunctions, and particles obeyed Wackernagel's law, which required unaccented clitics to occupy the second position of a clause, immediately following the first accented word. Common enclitic conjunctions included *-kʷe ('and') and *-wē ('or'), which attached to the second coordinated word. Subordinate clauses typically preceded main clauses, and relative clauses used the relative pronoun *yos (or *kʷos) placed near the beginning of the clause. Prepositions were originally independent spatial adverbs or postpositions that subsequently fused with verb roots as preverbs or governed specific nominal cases.
Typological scholars have investigated the earlier morphosyntactic alignment of the proto-language. Christianus Uhlenbeck in 1901 and Georgiy Klimov proposed that the historical nominative-accusative alignment was preceded by an active-stative or ergative alignment. Evidence cited in support includes the identity of the nominative singular of neuter nouns with the accusative (reflecting an unmarked absolutive case), the distinctive *-s marker restricted to the nominative singular of animate nouns (reflecting an ergative marker), and the division of verbal morphology into active and stative categories. Roland Pooth proposed a templatic model interpreting roots as consonantal frames superimposing vocalic patterns with a direct-inverse transitivity alignment.
Lexicon and culture
Linguistic paleontology uses reconstructed vocabulary to reconstruct the material culture, social organization, environment, and beliefs of Proto-Indo-European speakers. Over 1,200 secure lexical roots have been reconstructed. Kinship terminology was elaborate, and specific terms existed for relatives through the husband, including *sweḱuros ('husband's father'), *sweḱrúh₂ ('husband's mother'), *dayh₂wḗr ('husband's brother'), *ǵ(e)mHōr ('daughter's husband'), and *snusós ('son's wife'), whereas corresponding terms for relatives through the wife were largely absent; from this, scholars suppose that the society was probably patriarchal and that women moved into their husband's family after marriage.[1][2]
The pastoral and agricultural economy of the speakers is reflected in a rich agricultural lexicon. Domesticated animals included cattle (*gʷōws), horses (*h₁eḱwos), sheep (*h₂ówis), goats (*diks), pigs (*suHs), and dogs (*ḱwṓn). Wool (*Hwlh₁neh₂) was sheared from sheep and woven (*h₁webʰ-) into textiles. Agricultural terms include the plow (*h₂erh₃-trom), to plow (*h₂erh₃-), to sow (*seh₁-), field (*h₂eǵros), grain (*ǵr̥h₂-nó-), and quern-stone (*gʷreh₂uōn). The diet included meat (*mēms), milk (*h₂melǵ-), butter (*h₃(e)ngʷ-n̥), cheese (*tuHris), salt (*seh₂l-), honey (*melit), and mead (*medʰu).
A body of vocabulary relates to wheeled transport, including wagon (*weǵʰnos), axle (*h₂eḱs-), yoke (*yugóm), and three distinct words for wheel: *kʷekʷlóm (a reduplicated form of *kʷel- 'to turn'), *Hroth₂os (from *Hret- 'to run'), and *Hwr̥gi-. Wheels were made of three joined planks cut into a circle.[3][4] Because the daughter languages share words for wheel and axle, most researchers place the split of the language no earlier than 3400 BCE, the date to which archaeology assigns the first secure use of wheels, including in the assumed language area.[5] Spoked wheels appeared around 2500 to 2000 BCE, after the breakup of Proto-Indo-European.[3][4]
Environmental vocabulary encompasses fauna such as the wolf (*wĺ̥kʷos), bear (*h₂ŕ̥tḱos), fox (*wl(o)p-), lynx (*luḱ-), beaver (*bʰébʰrus), otter (*udros), deer (*h₁elh₁ḗn), elk (*h₁ólḱis), and aurochs (*tauros). Flora includes birch (*bʰerHǵos), oak (*pérkʷus), beech (*bʰeh₂ǵos), ash, maple (*h₂ēkr̥), willow (*weit-), and yew (*taksos). Reconstructed religion centered on a sky father deity, *Dyḗus ph₂tḗr ('Daylight-Sky Father'), paired with *Dʰéǵʰōm méh₂tēr ('Earth Mother'), the solar deity *Seh₂ul, the dawn goddess *H₂éwsōs, the Divine Twins (*Diwós sūnú, 'Sons of Dyēus'), and a storm god associated with the oak (*Perkʷunos). Poetic formulas survived across branches, including *ḱléwos n̥dʰgʷʰitom ('undying fame'), preserved in Greek kleos aphthiton and Vedic sravas aksitam.
Homeland and archaeogenetics
The location of the Proto-Indo-European homeland (Urheimat) has been debated since the nineteenth century. Early proposals suggested northern Europe, Scandinavia, Central Europe, or the Indian subcontinent. In 1886, Otto Schrader proposed the Pontic-Caspian steppe. In 1956, Marija Gimbutas synthesized archaeological data into the Kurgan hypothesis, proposing that the pastoralist Yamnaya culture and related kurgan (burial mound) cultures in the Pontic-Caspian steppe north of the Black and Caspian Seas between 4500 and 2500 BCE represented the speakers of Proto-Indo-European. The domestication of the horse and the adoption of wheeled wagons facilitated pastoralist expansions across Eurasia in successive waves.
The leading alternative theory was the Anatolian hypothesis, formulated in 1987 by archaeologist Colin Renfrew. Renfrew proposed that Proto-Indo-European originated in central Anatolia (around sites like Catalhoyuk) in the seventh to sixth millennia BCE and expanded into Europe alongside the spread of early agriculture. Other hypotheses include the Armenian hypothesis, formulated in 1984 by Tamaz Gamkrelidze and Vyacheslav Ivanov, which placed the homeland in the Armenian Highlands south of the Caucasus; the Central European or Balkan hypothesis, associating the language with the Linear Pottery culture; and the broad homeland hypothesis, viewing most of Europe as a prehistoric continuum.
Beginning in 2015, archaeogenetic studies of ancient DNA by Wolfgang Haak, Iosif Lazaridis, David Reich, and others lent support to the steppe hypothesis.[6][7] Genome-wide analyses revealed that the Yamnaya steppe pastoralists contributed substantial ancestry to the Late Neolithic Corded Ware culture of Central Europe, amounting to roughly seventy-five percent of their genetic profile. Following these discoveries, Colin Renfrew accepted the reality of migrations of populations speaking one or several Indo-European languages from the Pontic steppe towards Northwestern Europe.[8][9] A 2025 genetic study by Lazaridis and colleagues further traced the formation of steppe pastoralists to a Caucasus-Lower Volga genetic cline, with some scholars arguing for an ultimate pre-Yamnaya origin south of the Caucasus for Proto-Indo-Anatolian.
Macrofamily hypotheses
Linguists have investigated whether Proto-Indo-European shares deep genetic relationships with other language families. The most widely studied hypothesis is Indo-Uralic, which posits a common ancestor for Proto-Indo-European and Proto-Uralic. Scholars such as Vilhelm Thomsen, Karl Bernhard Wiklund, Bjorn Collinder, Frederik Kortlandt, and Alwin Kloekhorst have pointed to structural parallels in pronominal roots (first-person *m-, second-person *t-), nominal case affixes (accusative *-m), and verbal inflection. Indo-Europeanists assess these proposals differently: Robert Beekes considered it justified to regard the Indo-European and Uralic families as related, Michael Meier-Brugger held that a relationship between Indo-European and other families can be neither proven nor disproven, and Herzenberg judged even the Indo-European and Uralic correspondences insufficient for a full comparative grammar.[10]
The Nostratic hypothesis, proposed by Holger Pedersen in 1903 and developed in the 1960s by Vladislav Illich-Svitych and Aron Dolgopolsky, groups Indo-European with Uralic, Altaic, Dravidian, Kartvelian, and Afroasiatic into a macrofamily. Joseph Greenberg proposed a related Eurasiatic macrofamily, while Harold Fleming and Nikolai Andreev formulated broader Borean constructs. John Colarusso proposed a Proto-Pontic connection between Indo-European and Northwest Caucasian, though Johanna Nichols and others noted that their morphosyntactic structures differ fundamentally. Many Indo-Europeanists and comparative linguists do not accept the Nostratic hypothesis and regard it as unconvincing at best and wrong at worst, and many linguists reject the method of mass lexical comparison on which such proposals often rest.[11][12][13]
Reconstructed texts
Because Proto-Indo-European is an unattested proto-language, scholars have composed illustrative texts to demonstrate its reconstructed grammar and phonology. In 1868, August Schleicher composed the fable The Sheep and the Horses (Avis akvasas ka). Schleicher's original text reflected a Sanskrit-based phonological model lacking the e/o ablaut distinction and laryngeals. As Indo-European linguistics advanced, the fable was repeatedly revised, notably by Hermann Hirt in 1939, Winfred Lehmann and Ladislav Zgusta in 1979, and Andrew Miles Byrd in 2013, whose version incorporated comprehensive laryngeal and syllabification reconstructions.
In the 1990s, Sukumar Sen, Eric Hamp, and others composed a second illustrative text titled The King and the God (rēḱs deiwos-kʷe), based on a narrative from the Rigveda in which a childless king prays to the god Varuna for a son. Reconstructions of these compositions are recognized by linguists as educated approximations; in 1969, Calvert Watkins remarked that despite more than a century of comparative scholarship, linguists could not definitively guarantee the exact natural phrasing of a single complete Proto-Indo-European sentence.
In popular culture
Proto-Indo-European has appeared in several modern media works. In Ridley Scott's 2012 science fiction film Prometheus, the android character David studies historical linguistics during interstellar flight, reciting lines from Schleicher's fable, and later addresses an extraterrestrial Engineer using spoken reconstructed Proto-Indo-European. In 2014, American composer Christopher Tin composed Water Prelude, the opening movement of his classical crossover choral album The Drop That Contained the Sea, featuring lyrics sung in reconstructed Proto-Indo-European accompanied by the Royal Philharmonic Orchestra. In the 2016 Stone Age action-adventure video game Far Cry Primal, linguists developed three functional prehistoric dialects (Wenja, Udam, and Izila) based directly on Proto-Indo-European grammar and vocabulary.
Where editions disagree (3)
- English: The existence of *a as an independent phoneme is debated, with most instances explained by adjacent laryngeal *h₂.
- Lithuanian: A small number of roots cannot be explained by laryngeal theory, requiring the reconstruction of an independent vowel *a.
- Finnish: Comparative evidence for an independent *a phoneme separate from laryngeal *h₂ is considered insufficient by some linguists like Tijmen Pronk.
- English: Broad consensus reconstructs default SOV word order, though SVO was proposed by Jacob Wackernagel and Paul Friedrich argued for a VO ancestor.
- German: Winfred Lehmann reconstructed default SOV order based on typological features.
- French: Word order has been hypothesized to have shifted from SOV to SVO during the late stage as Anatolian split off.
- English: The Kurgan hypothesis places the homeland in the Pontic-Caspian steppe, debated against the Anatolian hypothesis.
- Russian: The Kurgan steppe hypothesis, Anatolian hypothesis, Armenian Highland hypothesis, and Balkan hypothesis remain competing theories.
- Portuguese: Ancient DNA studies in 2015 strongly supported the steppe homeland over Anatolia.
Sources (89 Wikipedia editions)
The Lithuanian, Polish, Russian, and German editions provide extensive tables and grammatical treatments absent from the English article, including complete inflectional paradigm reconstructions for nouns, pronouns, and verbs across four distinct scholars (Sihler, Beekes, Fortson, and Ringe). They also detail the historiography of early pioneers like Boxhorn, Coeurdoux, and Andreev, analyze the morphological mechanics of the Caland system and Wackernagel's law, and provide explicit transcriptions of Schleicher's fable across various historical revisions. Furthermore, Russian and Lithuanian sources present comprehensive lists of reconstructed fauna, flora, kinship, and agricultural vocabulary, along with extensive coverage of macrofamily hypotheses such as Indo-Uralic, Nostratic, and Borean.
Assembled from the Wikipedia articles below, each pinned to the revision read on 2026-09-27. Together they hold 1864 references; the English article alone has 76.
References
- Beekes, R. (2011). Comparative Indo-European linguistics: an introduction (2 leid.). Amsterdam — Philadelphia: John Benjamin’s Publishing Company. pp. 39. ISBN 978-9-02-721186-6.
- Beekes, p. 39
- Mallory, Adams, 2006, p. 247—249. (Adams D. Q., Mallory J. P. The Oxford Introduction To Proto-Indo-European And Indo-European World. — Oxford: University Press, 2006.)
- Beekes R. S. P. Comparative Indo-European linguistics: an introduction. — 2 ed. — Amsterdam — Philadelphia: John Benjamin’s Publishing Company, 2011. — P. 38. — ISBN 978-9-02-721186-6.
- Fortson, 2.58f
- Haak; et al. (2015). «Migração em massa da estepe é fonte das línguas indo-europeias na Europa» (pdf) (em inglês). 2015. 172 páginas. Consultado em 6 de novembro de 2015
- Haak, Wolfgang ym.: ”Massive migration from the steppe was a source for Indo-European languages in Europe”. Nature, 2015, 522, s. 207–211. doi:10.1038/nature14317.
- Renfrew, Colin (8 November 2017). Marija Redivia : DNA and Indo-European origins. Chicago: Marija Gimbutas memorial lecture – via YouTube.
- Pellard, Thomas; Sagart, Laurent; Jacques, Guillaume (2018). "L'indo-européen n'est pas un mythe". Bulletin de la Société de Linguistique de Paris (in French). 113 (1): 79–102. doi:10.2143/BSL.113.1.3285465. S2CID 171874630.
- Красухин К. Г. (2013). „Новые руководства по индоевропейскому языкознанию“. 6. Вопросы языкознания: 115, 132. ISSN 0373-658X.
- George Starostin. Nostratic . Oxford Bibliographies. Oxford University Press (29 октября 2013). doi:10.1093/OBO/9780199772810-0156. — «Nevertheless, this evidence is also regarded by many specialists as insufficient to satisfy the criteria generally required for demonstrating genetic relationship, and the theory remains highly controversial among mainstream historical linguists, who tend to view it as, at worst, completely invalid or, at best, inconclusive.» Архивировано 13 сентября 2015 года.
- Clackson J. Indo-European Linguistics. — Cambridge: Cambridge University Press, 2007. — P. 20. — (Cambridge Textbooks in Linguistics). — ISBN 0-52-165367-3. — [Архивировано 4 марта 2016 года.] — «The frustration evident in many of the statements of Nostraticists is clear: they are using the same methods as IE linguists, yet their results are not accepted by most IE linguists for reasons which are seldom clearly articulated.»
- Lyle R. Campbell: Beyond the Comparative Method?. W: Barry J. Blake, Kate Burridge, Jo Taylor: Historical Linguistics 2001. Amsterdam: John Benjamins Publishing, 2003, s. 35-39. DOI: 10.1075/cilt.237.05cam. ISBN 90-272-4749-8. (ang.).
