A distributional broadsheet Indus Valley Civilisation
2,543 artefacts
11,135 sign tokens
515 identifiable signs
41 rounds of testing

Indus Valley script analysis

Forty-one hypotheses about the Indus script, tested against the corpus itself. The controls killed most of them; what’s left is small, and it holds.

This is not a decipherment

What survived

Eleven results that cleared every control

Every one of these is a distribution, never a meaning. Position, site and object class stay fixed inside each null, duplicated texts get collapsed before anything is counted, and each headline was re-run 300 times through a noise model built from the disagreement between two independent digitizers.

Texts end in a fixed cohort

Seven signs close 1,179 of 2,086 texts where a within-text position shuffle predicts 400. Signs 740 and 520 turn up together six times against 27.7 expected, and that pair is the one exclusion still standing after every control we have.

finality z = +47.5 · pair z = −5.45

The ending floats, but only one step

The construction sits at −2 whenever 400 or 90 follows it, so the paradigm floats rather than being nailed to the last position. Behind it there is nothing: 33 candidate exclusions at −2 and not one survives.

−2 cohort: 0 of 33 exclusions after BH

A text never reuses a sign

Repetition sits far below what the exact-position, site- and object-matched control predicts. It is the most reliable fact in the corpus, and it explains why positional statistics keep working here while sequence statistics dissolve.

z = −14.85 · noise 95% −14.17 to −10.03

Seals and tablets fill the last field differently

Both media draw on one inventory and one layout, then part company at the end. Tablets append 400 where seals mostly don’t, and seals carry numerals nearly twice as often.

p = 1e−14

One form across the civilisation

Seal endings at Mohenjo-daro and Harappa can’t be told apart, and the smaller towns follow the same template. Harappa’s tablets are the only local practice in the corpus.

p = 0.98

Stroke signs are numerals, and twelve is a spike

Counted off the rendered glyphs rather than inferred from database ids, the values run 1 to 9 with an ordinary decay and then jump at twelve. Sign 55 carries 37 deduplicated tokens; no decay curve reaches that.

Poisson p = 1.4e−45

Numeral side is sign-specific

Pooled across the corpus the order comes out nearly 1:1, which hid the structure completely. Fifteen of 37 eligible signs are at least 90% one-sided against 0.01 expected, and the overdispersion holds through every control and every noise draw.

controlled z = +16.09

Dependence reaches distance four

Against a surrogate that preserves every observed bigram exactly, mutual information stays above null at separations of two, three and four signs. Five is borderline and six is nothing. Most of what the corpus has is local.

excess 0.052 to 0.078 bits, BH corrected

The four headlines replicate

Repeat avoidance, the 740/520 exclusion, terminal concentration and numeral-side overdispersion all hold in random halves and in disjoint site halves, where no site appears on both sides. They also sit outside nulls built by destroying the pairing the pipeline looks for.

exploratory paradigm scan FPR = 7.1%

The inventory is not saturated

515 identifiable types once the twelve database markers for unidentified signs come out. Estimators that key on singletons put the real total near 700, and the accumulation curve hasn’t turned over yet.

Chao1 ≈ 715 · ACE ≈ 695

Copies are mostly local production

250 sequence types recur, producing 684 attestations beyond the first, and 84.4% of those are another copy where the text already sat. Circulation isn’t nothing though: 35% of repeated types reach two to four sites.

577 local · 107 new-site attestations

House rules

Six rules, five of them adopted after something went wrong

Most of these came out of a mistake, and each one has a body count. They sit in the order the mistakes happened.

01

Epigraphy only

The source database ships a Sanskrit decipherment that its own field rejects. We took the epigraphic layer and nothing else: which signs, in what order, on which object.

02

Deduplicate first

Mass-produced tablets repeat one text dozens of times, and left in they manufacture significance. This rule cost 25 sign–motif findings on the day it went in.

03

Control for position

If one sign prefers the start and another prefers the end, ordering comes free. Every null here shuffles inside exact positions.

04

Control for site and object class

Two signs from different cities never meet, for reasons that have nothing to do with grammar. In this corpus medium and city are very nearly the same variable.

05

Define groups before you look at the outcome

Sorting signs by the behaviour you’re about to test guarantees a result and teaches you nothing. Round 21 withdrew a shape split for exactly this.

06

Report the failures

Everything below is the record, including the rounds that knocked over earlier headline results of our own.

The ledger

Everything we tried, in the order it happened

Each round came out of the one before it. When a round overturns something we had already concluded it gets tagged Corrected and the original stays on the record, instead of being quietly edited away. Open a round for the evidence and its plate.

  • Held under every control
  • Narrowed to something smaller
  • Hypothesis refuted
  • Corrected an earlier claim of ours
  • Blocked by data or method
  • Groundwork
Each square is one round, coloured by what happened to the hypothesis. Click a square to open that round below.

The wall

Why this stops where it stops

The corrected merged corpus records 527 sign IDs, twelve of which are database markers for unidentified signs, leaving 515 identifiable types. Roughly 416 non-numeral signs still have fewer than 20 tokens. Partial pooling puts all of them on one uncertainty scale but it can’t invent observations: three rare signs came out with finality intervals above the base rate, two of those look like post-terminal additions, and the model fails its global posterior-predictive check.

515 isn’t a ceiling either. Estimators sensitive to singletons put the merged total near 695–715 and the accumulation curve hasn’t plateaued, so more data will add types as well as tokens. That makes the wall bigger rather than smaller.

Two independent digitizers agree on 93.2% of aligned signs. Propagate that disagreement and the no-repeat effect, the terminal cohort, the 740/520 anchor and the global numeral-side result all come through intact, while most marginal pairwise edges get erased. The wall is explicit uncertainty now rather than a blanket objection: strong effects clear it, and z around −2 to −3 generally doesn’t.

ICIT (4,660 artefacts, 17,957 signs) is the only corpus big enough to move any of this, and it’s request-gated. Round 41 showed our numbering is already ICIT’s, so integrating it would need no crosswalk. The gain would be stratification rather than raw power; doubling N buys about √2 on a z-score and leaves the median sign just as rare.

The corpus, measured
Artefacts in the parsed corpus2,543
Sign tokens11,135
Records after deduplication2,086
Recorded sign IDs (merged)527
Identifiable types515
Non-numerals under 20 tokens416
Estimated true inventory (Chao1)~715
Inter-digitizer agreement93.2%
Tokens perturbed per noise draw668 (7.3%)
ICIT, if it can be obtained4,660 / 17,957