On Sandhi in Pali by the late R. C. Childers

Originally Published by Journal of the Royal Asiatic Society,
Volume 11, Issue 1, January 1879, pp. 99–121.

revised with additions and corrections by Ānandajoti Bhikkhu

From the introduction:

The rules of sandhi in Pāḷi are the effect of the spoken word and a living, spoken—not literary—language of its time. All Prakrits have similar features: unpredictability, and a confusion of mights and maybes. Childers’ work on the subject was made when only a small number of works had been published, when rules for presentation were still in flux, and when even the way we represent Indian words and syllables had not been standardised. It is still to this day the only work I know of its kind in the grammatrical literature.

From the General Remarks section of the book:

In Sanskrit sandhi is imperative, in Pāḷi it is to a great extent optional: between separate words it takes place but seldom, and even in compounds hiatus occurs. Again, while sandhi is regular and uniform in Sanskrit, in Pāḷi it is very irregular. For example, while in Sanskrit na upeti must always become nopeti, in Pāḷi it might become nopeti, or nupeti, or nūpeti, or remain na upeti without sandhi change taking place.

11 Likes

It’s one of my favourite papers on Sandhi in Pāli. Great that Ven. Ānandajoti has rejuvenated it.

3 Likes

I find this study very interesting in a linguistic sense. It is a very educative compilation (thanks to Childers, Anandajoti and yourself for sharing), the contents of this work need to be discussed here in great detail as the implications of Sandhi in Pali have a massive lot to inform us about the history and origins of Pali itself if we consider their logical implications properly, but Childers is partly wrong about his claims regarding both Sanskrit sandhi and Pali Sandhi.

Firstly Sandhi is not universal or mandatory in Sanskrit. The picture is far more nuanced.

There are two kinds of Sandhi: external and internal

  1. Sandhi between words i.e. external - This is optional in normal written prose/speech and mandatory in literary creations (both prose/gadya or verse/padya)
  2. Sandhi inside words i.e. internal - This is mandatory in all contexts.
  3. In Prose, Sandhi although optional between words is usually applied by most classical writers in written texts. In speech it is truly optional and inconsistently applied.
  4. The option to write unsandhied Sanskrit (external sandhi) is usually exercised only by, or for the benefit of people who would struggle with sandhied Sanskrit i.e. for the benefit of students and those who are not familiar with Sanskrit as it keeps word boundaries and word forms clear and simple. Most classical writers never exercise the option to write unsandhied sanskrit i.e. they always write with sandhi.
  5. In speech even those who do regularly apply external sandhi in written prose dont apply them uniformly when speaking.
  6. There are some words and phrases and contexts in which are spoken only with sandhi even when speaking.

Vedic texts too follow mostly (but not 100%) the same sandhi rules as classical sanskrit, Vedic word forms and classical word forms have a large overlap as well. So sandhi-wise classical sanskrit simply continues Vedic sandhi rules for the most part.

About Pali Sandhi a whole lot needs to be said - but Pali Sandhi should be studied historically by comparing it with Ardhamāgadhī, Ashokan Epigraphic language, and Classical and Vedic sanskrit to make full sense of why Pali uses some divergent sandhi norms (the Pali/Ardhamāgadhī/Aśokan sandhi rules that are common to classical and Vedic sanskrit can be considered normal).

If anyone wants to join me in analysing this text (can we do this here?), pls let me know.

4 Likes

I suppose it would be a fine thread as any, even better if the discussion could somehow relate to / build on the aforementioned work — where do the ancient and modern scholarship agree and challenge with the findings, etc.? :slight_smile:

Though plebs like me can only really watch from the sidelines! :sweat_smile:

3 Likes

Fine with me. Seems useful.

4 Likes

Anandajoti says in the introduction:

The rules of sandhi in Pāḷi are the effect of the spoken word and a living, spoken—not literary—language of its time. All Prakrits have similar features: unpredictability, and a confusion of mights and maybes.

This understanding is curiously in direct contradiction with the evidence presented by Childers.

Childers posits Sanskritic grammar and word-forms underlying surface-level Pali word forms, as in the following observations:

It would not only be a misapplication of labour, but positively misleading, to work out rules of internal Sandhi from, for instance, such forms as sabbhi and lacchate. Our only proper course is to trace them to their Sanskrit originals sadbhis and lapsyate, and bring them under rules, not of Sandhi, but of phonetic change.

In Sanskrit, the rules of sandhi for words and for compounds are the same; in Pāḷi, they present several points of difference. Pāḷi compound words are of two classes—first, compounds, which are phonetic corruptions of corresponding Sanskrit compounds; and, secondly, compounds in which two Pāḷi words are independently combined, without reference to Sanskrit. To the first my remarks at p. 100 on internal sandhi are applicable; they are by their nature excluded from the department of sandhi. Thus jaraggava cannot be brought under any Pāḷi rule of sandhi; all we can do is to trace it to an older Sanskrit form jaradgava, of which it is a phonetic corruption. On the other hand, kulitthi, at Par. 3, cannot be identified with a Sanskrit compound kulastrī, but is an independent combination of the Pāḷi form itthi with kula.

Kula & strī are common words (and once stri is phonetically transformed to itthi, it is impossible to then follow the same sandhi rules as in Sanskrit as the underlying phonetic form of the word has changed). So if Childers is expecting kula & itthi not to be compounded independently in Pali after the underlying members have changed phonetic form, I thiink his expectation does not make sense.

Childers also draws a hierarchy of Sandhi application in Pali. He says

In verse, word-sandhi is much less restricted and much more frequent than in prose, being in great measure governed by the question of metrical exigency. Thus in the first two pages of Dhammapada there are nine sandhis, of which only two, nādhigacchanti and nappasahati, would occur in prose. The remainder, for instance sammantīdha, vantakāsārassa, are used metri causā. Some of the bolder sandhis, as the elision of and aṁ, are confined to verse.

Sandhi is more extensively used in the early texts of the Tipiṭaka than in the late texts of the Commentaries.

So given that there are at least 4 forms of Pali based on incidence of Sandhi usage (Verses, Late Verses, Prose and late Prose), which of these does Anandajoti consider closest to his putative “living language”? Also if Pali was a living language rather than a literary register, how does Childers say that there are words which only make sense if interpreted as phonetic corruptions of Sanskrit words?

Further Childers mentions

ati and paṭi before a vowel generally become acc and pacc, standing for an older aty and paty. Ex. accuṇha=ati-uṇha, accokkaṭṭha=ati-okkaṭṭha, accodāta=ati-odāta, accagā=ati-agā, paccāroceti=paṭi-āroceti, paccaṅga=paṭi-aṅga, paccupaṭṭhita=paṭi-upaṭṭhita, paccaññāsi=paṭi-aññāsi.

The older aty & paty (sic) mentioned above are Sanskrit (paty should rather be praty).

  1. In one case followed by a vowel becomes jj, which represents an older dy: najjantara=nadī-antara.

This older dy is also Sanskrit.

  1. abhi and adhi before a vowel generally become abbh and ajjh, which represent older forms abhy and adhy. Ex. abbhaññāsi=abhi-aññāsi, abbhattha=abhi-attha, abbhokāsa=abhi-okāsa, ajjhabhāsi=adhi-abhāsi, ajjhāvasatha=adhi-āvasatha, ajjhokāsa=adhi-okāsa, bojjhaṅga=bodhi-aṅga.

The older abhy and adhy are Sanskrit again

In one instance, ativiya=ati-iva, we have y inserted between two is, for ativiya points to a transition form atiyiva.

Here this explains how the earlier iva (Sanskrit) becomes Pali viya

  1. Occasionally when a word ending in a vowel is compounded with a word beginning with a consonant, a consonant which originally belonged to the base of the first word is revived, and if necessary assimilated to the initial consonant of the second word. Thus sammā-paññā becomes sammappaññā, which probably represents the Sanskrit samyakprajñā; anto is the Sanskrit antar, but in composition we sometimes have the original r revived, e.g. antaraghara=Sanskrit antargṛha. Again, the Sanskrit base catur is catu in Pāḷi, e.g. catuvagga (catumāsaṁ), but compounds like catugguṇa, catubbagga, catummukha, point to Sanskrit forms caturguṇa, caturvarga, caturmukha retaining the final r. So also we have cha=Sansk. ṣaḍ, but chammāsa points to an original ṣaḍmāsa. And puna compounded with bhava and with puna gives punabbhava and punappuna, in Sansk. punarbhava and punaḥpunar. [120]

More examples of Pali Sandhi transformations of underlying Sanskrit vocab.

What all the above shows is that Prose Pali is a phonetically standardized register, and the Commentarial Prose is even more standardized and artificial than the prose of the Canonical Pali. The evidence also shows that natural phonetic principles do not apply to Pali sandhi (which would be expected if Pali itself were a spoken natural dialect/language), and these rules were idiosyncratically applied by the composers of the suttas. However the language that underlies Pali (rather than Pāli itself) was a natural spoken language, and that language points to Sanskrit or something very close to it.

3 Likes

Thank you for the interesting remarks. :slight_smile:

Isn’t it possible that Pāli represents not just one, but several Indic languages and/or dialects in close proximity, all bubdled up together organically from a shared bur ultimately different communities?

And for example, French is notorious for having a haywire grammar and spelling, applied haphazardly. We would be amiss to consider French an artificial language though, yes?

There is no evidence at all for any such multiple Indic languages of which Pāli could have been a mixture. No such languages are recognized or attested by anyone from the historical or literary evidence - so there is no reasonable way to assume such a thing. However there is evidence to the contrary, that it was based on the underlying speech of one region, Western-India.

Pali (and perhaps Gandhārī too) is what the Indian native tradition named ‘Paiśācī’ (paiśācī means ‘something that belongs to piśāca-s’, and piśāca means ghost, ghoul or demon) i.e. the language was associated with ghosts from the the middle of the first-millenium CE (the time of the Pali commentaries onwards). The native Indian tradition remembers just one lost work in this language - called Bṛhatkathā (a work supposedly as vast or bigger than the Pali Canon) but of which not even one or two verses survive in the present day.

A very good recent paper about Paiśācī - Pāli relations can be found at http://prakrit.info/resources/papers/2014-ghosts_from_the_past-ollett.pdf

Athough the linguistic properties of Paiśācī as described in other later literature appear to coincide with Pāli for the most part, no scholar has satisfactorily explained how this language has come to be associated with piśācas (ghosts/ghouls). My own interpretation of this name is that the name Paiśācī is a hyper-sanskritization (incorrect sanskritization) of the Middle-Indic/Pali word pesacci, which is actually to be traced to the Sanskrit pāścātyā (meaning “western”, i.e. a ‘western language’). Pāli is already known to be derived from a western language associated with the Western Indian state of Gujarat. Gandhārī, with which Pāli shares characteristic Middle-Indic features, is also a western-language (of the north-west i.e. Gandhāra and Punjab). But since Pāli / Paiśācī is a literary language only, it had no life as such in speech and hence there is no literature except the Pāli canon and the lost work composed by Guṇāḍhya (the Bṛhatkathā).

3 Likes

I would think that multiple grammatical forms for the same word would point toward something to that effect (i.e. Pali canon swallowing local dialects / languages).

I think it would be a reasonable inference that parts of that region spoke sibling languages / dialects which would be swallowed by the Magadhi / Pali as Buddhism got more widespread, no? They would not need to leave behind any literature, as most languages of the world leave practically no trace at all (except the aforementioned different grammatical forms attested in Pali for the same word).

Again, just a wild inference. :slight_smile:

All of this said befire I had a chance to check this out:

Thanks for that, should be interesting. :slight_smile:

Do you suppose Pāli was something like an artificial koine for Western dialects that underwent some Sanskritisation? Or something else?

The grammatical forms in Pali are generally fairly regular, it is the phonetic forms that indicate widespread confusion, and that confusion is not due to dialectal diversity.

All the examples of dialectal vocabulary given in the Pali canon only indicate the existence of multiple words referring to the same/similar thing, not multiple phonetic forms of the same word.

For example in MN139 examples of janapada-nirutti (regional vocabulary) referring to the same/similar object (a saucer) is mentioned -
“Idha, bhikkhave, tadevekaccesu janapadesu ‘pātī’ti sañjānanti, ‘pattan’ti sañjānanti, ‘vittan’ti sañjānanti, ‘sarāvan’ti sañjānanti ‘dhāropan’ti sañjānanti, ‘poṇan’ti sañjānanti, ‘pisīlavan’ti sañjānanti.”

Here pātī (Sanskrit pātrī), pattam (Sanskrit pātram), vittam (?), sarāvam (Sanskrit śarāvam), dharopam (?), poṇam (Sanskrit pāna), pisīlavam (Sanskrit piśīlava) etc are not an example of Pali borrowing words from multiple Indo-Aryan languages but inheriting regional vocabulary that already also existed in Sanskrit.

So the multiple phonetic/lexical variants of the same words that you’re talking about is not dialectal but simply a result of phonetic confusion.

Koine (meaning “common” in Greek) was the standardized, widely spoken form of Greek that served as the lingua franca across the Eastern Mediterranean from the 4th century B.C.E. to the 6th century C.E.. Emerging from Alexander the Great’s conquests, it merged Attic and other dialects into a simplified, accessible language used in trade, politics, and the New Testament.

If this is what you mean by the word koine, Pali doesn’t appear to be a koine. A koine would not belong to just Buddhism in the way Pali is almost exclusively a Buddhist canonical literary register.

It appears to be a literary-only (not spoken) register that was unique to early-Buddhist Texts. Thus it was not a lingua franca used outside Buddhism and could not have been a koine.

The Pali canon contains so many templated phrases where so many different people utter exactly the same words in similar situations that to me looks like the Pali suttas were put together like lego pieces consciously constructed to accord with specific templated doctrinal models. They dont seem to be what people actually spoke but what they were expected to speak.

Many of the people who supposedly spoke with the Buddha would not historically have spoken Pali. I find even the idea that the Buddha spoke Pali very doubtful.

Apparently there must have been a core canon written in a pre-Pali natural language of which the Pali canon is a greatly expanded and heavily edited and standardized successor.

Phonetically there doesnt appear to be any historically attested language that matches Pali from the time it supposedly originates. It doesnt appear to have evolved naturally from a prior Indo-Aryan dialect or evolve further into a later Indo-Aryan dialect.

If Magadha is the area that is currently recognized to be (i.e. the area in and around the Indian state of Bihar) then Pali is not the language of Magadha for sure and therefore cannot be called Māgadhī. Else if it was actually based on a Magadhan language, Magadha would have to be a western kingdom close to the Indian state of Gujarat, and Kosala (where Śrāvastī/Sāvatthī was located) would have to correspond to somewhere around the Indian state of Punjab. So the geographical provenance of Pali is also apparently misdiagnosed.

So Pali is for Buddhism (in my understanding) what Ecclesiastical Latin is for Roman Catholicism, except that Pali is even less of a natural language than Ecclesiastical Latin

3 Likes

Okay. So, as far as I understand, you believe that Pāli was specifically developed as a literary language, probably somewhere in Western India, right? Why, then? What for? Why did they use Sanskritized forms (Brahmana, -tva, etc.) instead of the expected phonological outcomes? Isn’t it possible that Pali had a predecessor in Western India and only became fossilized in Sri Lanka (cd. Russian Church Slavonic, which was never a spoken language as well)? Can’t stock phrases in Suttas be accounted for by regularisation for oral transmission of the texts?

For the record, I don’t have any stakes in this game, I don’t have any opinion on whether your theory is correct or not, just asking questions :slight_smile:

1 Like

I follow this. At the same time, the lego pieces derive from something (versus magically appearing). Referring to your Ecclesiastical Latin analogy, that depends linguistically on what were, originally, the sayings of Christ passed on orally/phonetically.

If the analogy is valid, aramaic has absolutely nothing to do with Ecclesiastical latin (or any latin for that matter). Obviously it doesn’t. Nonetheless aramaic > greek > formal latin does not negate the feasibility of a “Sayings practice” retaining its root infrastructure (the legos) throughout.

Am I pressing the “Sayings practice innovated by the Buddha” in your linguistically trained mind?

By the way, thanks so much for teasing out the sandhi discussion.

1 Like

I believe that the Pali Canon (talking only about the predominantly prose texts here) from the very beginning has been a written canon. There was no oral transmission of any prose text.

Since those texts (in a western dialect of Sanskrit) were originally written in the (Aramaic-derived) Kharoshthi script, which did not preserve complex syllables of Sanskrit very well, they were simplified in writing. Later these written texts were converted to Brahmi script which is when what looked like Gandhari (but was really malformed Sanskrit) came to be written with the double consonants etc that we now call Pali in the Brahmi script.

But the Brahmi script was capable of representing Sanskrit phonetically accurately, so other sects of Buddhists (the Sarvastivadins and Mulasarvastivadins) decided to convert it back into Sanskrit. However others like the Dharmaguptakas chose to retain it in Gandhari, while the Theravadins (or Sthaviravadins) maintained their canon in an artificial language that was between Gandhari on the one end and Sanskrit on the other end.

So I agree that Sanskrit Buddhist EBTS are sanskritized from a prior middle indic form, but thar middle Indic form itself originates from an earlier Sanskrit dialectal form.

In his paper Pali as an artificial language, Prof. Oskar von Hinüber deals with how Pali is artificial and an intermediate stage of converting those texts to Sanskrit, but he doesnt deal with why they were being converted to Sanskrit in the first place.

It is because as I say the Buddhists knew from their history that the Buddha must have spoken Sanskrit and since they were trying to move their texts from one script to another they were facing all these linguistic confusions and decisions of how to transform their texts without losing the meaning and authenticity.

Different sects of Buddhists could not agree about the form in which to retain their texts, but eventually everyone moved back to Sanskrit and Pali survives as an artificial language only because they moved geographically to Sri Lanka before their texts could fully be converted.

I dont see the connection between stock phrases and oral transmission. The Vedas had oral transmission without having any stock phrases. Stock phrases are neither a necessity nor an effect of oral transmission. Besides I don’t see any evidence of an oral transmission here.

Sure, I’ve been studying the historicity of the Pali language for years, and have not seen a theory that accounts for all the relevant facts. So I welcome all cross questions that help me test my hypotheses better.

2 Likes