Clarification on provenance and preservation of the Mahāsaṅgīti Pāli witness (snp2.4/pli/ms)

Dear SuttaCentral team,

Project Canon is an independent non-commercial research initiative focused on governance, provenance, and long-term traceability of canonical Buddhist source materials

We are currently evaluating the Pāli source for the Maṅgala Sutta (Suttanipāta 2.4).

From our investigation, we understand that the SuttaCentral page

snp2.4/pli/ms

corresponds to the repository file:

root/pli/ms/sutta/kn/snp/vagga2/snp2.4_root-pli-ms.json

in the bilara-data repository.

Before preserving this file as an unmodified research asset, we would appreciate clarification on a few points.

1. Does the witness code ms formally represent the Mahāsaṅgīti Tipiṭaka Buddhavasse 2500 edition?

2. Is the file under root/pli/ms considered the canonical Pāli witness corresponding to that edition?

3. What license or usage terms apply specifically to the Pāli root-text files under root/pli/ms?

4. May a research project preserve an unmodified snapshot of this file for long-term verification and provenance tracking?

5. Is there a preferred release, tag, commit, or archive package that should be cited or preserved instead of downloading individual files?

We are not requesting permission to modify the text, redistribute it, or create derivative editions.

Our goal is simply to preserve an authentic reference with complete provenance and traceability.

Thank you very much for your guidance.

Hello there. :slight_smile:

Yes to both, with an asterix.

For the details on the edition, you can view on a Pāli page:

Translation description

40 volumes of the Pali Tipiṭaka in Roman script. Includes the Nettiparaṇa, Peṭakopadesa, and Milindapañha. The Pali text on SuttaCentral was extracted from the source XML files of the Mahāsaṅgīti edition. Minor editorial changes have been made to numbering and punctuation, but no changes to the text itself.

Translation process

Prepared by the M.L. Maniratana Bunnag Dhamma Society Fund based on the Vipassanā Research Institute Tipiṭaka, 2537–2542, which was originally known as Chaṭṭha Saṅgāyana CD-ROM and is available at tipitaka.org. The entire text was proofread, and in addition, recited four times between B.E. 2543–2545 (2000–2002) and B.E. 2549–2550 (2006–2007) to verify every Pāḷi sound and to correct the printing errors of the original 40-volume edition. The primary reference for proofreading was the first edition of the Chaṭṭha Saṅgāyana Tipiṭaka published in Myanmar in 1958. Additional reference was made to 18 editions of the Pāḷi Tipiṭaka in various national scripts of other countries.

From the website:

  1. Public domain material

The original texts of Buddhism in Pali, Chinese, Sanskrit, Tibetan, and other languages are in the public domain. Such material does not fall within the scope of copyright.

Hi Shisanusha!

May I ask for some more details on this project? Who’s involved, what are the long term aims, etc. And do you have a website?

To further the asterisk as noted by Dogen, let me just give you the history.

I knew Major Surathat, leader of the Mahasangiti project. I visited their place in Bangkok and reviewed their work there, and in addition spent time with their team in Bodhgaya. When the Mahasangiti project collapsed, I was in touch and he assured me it would be back, but it never happened.

At that time we were linking to the Mahasangiti site for our Pali. So we needed a replacement. The best source we found was from Ven Yuttadhammo, who I’m sure you are familiar with. He had made a complete backup of the Mahasangiti files, and it is that which we use as the basis for our text.

You can find this here:

Use Git’s history to see any changes we made to this (mainly ṃ → ṁ).

If I have any questions about whether SC’s text is reliable, this is what I refer back to.

Any changes to SC’s Pali text consists of adjustments to:

  • punctuation
  • handling of sandhi in verse line breaks (Occasionally MS puts the sandhi consonant at the start of the second line, which is confusing)
  • normalizing spacing
  • markup (eg. paragraph breaks). But this is separate from the text.
  • headings (which are technically part of navigation not the text)

And similar things. By policy, we don’t change the Pali text at all, even in the (very rare) instances of obvious spelling mistakes. So if any changes are found, it’s by mistake.

The Mahasangiti team used a CC licence of some kind. However my understanding is that this has no legal force, as you cannot copyright an ancient text. Thus it applies only to their work (introductions, etc.)

All ancient text is inherently free of copyright, and there’s no legal restrictions on it.

I do ask that none of our work be used in any away for AI projects.

That would be the git archive linked above for the Mahasangiti.

Use this for our Pali text:

Note that this includes a couple of things that are not technically part of the MS edition, namely the dvemātikapāli, which should belong to VRI. I’ll move these files when I get the chance.