Baylee LaneBack to the essay

Companion notes · Revised October 2, 2026

Methods and further results

The definitions, comparisons, and limits behind What the Supreme Court’s questions reveal.

What the data covers

Oyez supplies case records, argument audio, speaker-labeled transcripts, and timestamps. The saved argument collection contains 1,408 arguments from October Terms 2005 to 2025. The article reports the results of that collection; it doesn’t claim that the latest term’s outcome records are complete.

The Supreme Court Database supplies ideological direction, majority membership, issue area, lower-court direction, and the winning side. Release 2025 Release 01 ends at October Term 2024. Its descriptive direction ledger contains 7,956 coded votes in divided cases. This is a larger sample than the cases that can be matched to usable argument measurements.

Petitioner votes are reconstructed from the coded winning side and majority or dissent membership. This binary representation simplifies mixed and multi-issue decisions. Ideological labels follow the Database’s codebook; they aren’t measures of motive. The argument model uses counts and timing, not the provisional legal-reasoning annotations in the explorer below.

Argument ledger, direction ledger, Oyez, and Supreme Court Database.

The question-counting comparison uses one common sample

The headline adaptation of Roberts’s rule counts substantive justice words addressed to each side. It predicts that the side receiving fewer words wins. A speaking turn is an uninterrupted contribution, so neither words nor turns are a literal count of individual questions. Rearguments are consolidated to one case outcome.

The common sample uses Oyez where its winner can be resolved and the Supreme Court Database where it can’t, within OT2005 to OT2024. Of 1,112 resolved cases, two have tied word counts and are omitted from both rules in the paired comparison. That leaves 1,110 cases:60.5% for word counting and 66.4% for always choosing the petitioner.

The paired 95% interval for the word-counting rule’s accuracy minus the petitioner rule’s is -9.9 to -1.9 percentage points. A saved sensitivity check that prefers SCDB where both sources resolve the case yields 60.5% for word counting. The headline retains its original, explicitly named source order.

Alternative ways of counting questioning
VersionCasesCorrectAccuracy
Words, all speakers111067260.5%
Words, principal advocates only111270062.9%
Words, excluding amicus sections74346162.0%
Speaking turns108162557.8%
Unanimous decisions only47530965.1%
Divided decisions only63536357.2%

Each alternative has its own usable case set. The unanimous and divided subsets, for example, can have different petitioner win rates. Their accuracies shouldn’t be compared with the full-sample baseline as though they were on one denominator.

The earlier 77% figure came from the Oyez-resolved subset. Cases dropped by that resolver had a lower petitioner win rate when matched to SCDB. The included, excluded, and combined samples, plus the term-level breakdown, remain in outcomeTestSelection. Missingness changed the apparent size of the benchmark advantage.

Correction: the ideological prediction test used an outcome-derived input

An earlier version reported that argument measurements predicted Gorsuch’s ideological votes more accurately than his voting history. That predictive interpretation is withdrawn. The model oriented its word imbalance using SCDB’s lower-court direction field. The codebook explains that this field is often constructed from the Supreme Court’s own ideological direction and whether it affirmed, reversed, or remanded.

The lower court’s actual ruling existed before argument. That doesn’t make a later field derived using the Supreme Court’s answer a valid pre-argument input. Training on earlier years can’t fix information about the test outcome entering a test feature. The earlier direction comparison, including its per-justice intervals, therefore can’t establish predictive performance before decision. A correction for testing multiple justices doesn’t address that problem either.

The original direction results remain archived. Their historical predictions were reproduced for the audit, including Gorsuch’s 185 correct votes out of 245. They aren’t evidence for the revised essay’s predictive claims. The descriptive vote map uses final vote direction as a description of outcomes, which is a separate use. The petitioner-vote models below exclude lower-court direction entirely.

The revised test adds argument to the same background model

All models are regularized logistic regressions. They produce a probability; at least one-half becomes a petitioner prediction for the accuracy score. The primary test covers 8,878 measured votes in 1,136 cases, OT2007 to OT2024. These cases represent 1,130 distinct SCDB decisions because some decisions cover related dockets.

The background model uses the justice’s previous petitioner-vote share, previous conservative-vote share, appointment party, source court, jurisdiction, and whether the United States is named as petitioner or respondent. It includes party-by-United-States interactions, previous petitioner-vote share for the justice within the source court, and historical petitioner case-win rates overall and by source court. Justice histories shrink toward the historical overall rate with ten prior observations; source-specific histories use twenty. Historical ideological share uses a symmetric ten-observation prior.

The United States indicator matches the normalized names “United States” or “United States of America.” It excludes separately named agencies, officials, and amicus appearances. Source court and jurisdiction are categorical, with categories learned only from training rows. Unseen categories receive zero on the learned category indicators. Missing party names have a separate indicator.

The argument inputs are word imbalance between sides, whole-bench imbalance, the difference between those imbalances, speaking share, speaking share relative to the justice’s earlier history, words per turn, and words per second. The combined model adds exactly these inputs to the background model. It doesn’t classify legal reasoning, tone, or persuasiveness.

The first training period is OT2005 to OT2006. Each later term is held out in sequence. Training outcomes must be dated before October 1 of the test year; every row’s history features use strictly earlier terms and decisions already available by that row’s term cutoff. Current-case ideological direction, lower-court direction, outcome, vote margin, and majority membership never enter the predictive features. Standardization uses training rows only.

The fixed primary L2 penalty is 10 on summed logistic loss, with an unpenalized intercept. Penalties 1 and 100 are reported sensitivities; the test results weren’t used to pick a winning penalty. All 360 fits converged. The saved model records include coefficients, feature names, and training means and scales.

These are retrospective records, not a timestamped reconstruction of every input as it existed before argument. The model also lacks briefs and independently coded ideological positions for both sides. It is a more detailed background comparison, not a comprehensive political or legal benchmark. Its accuracy is actually below the constant petitioner prediction, which is reported alongside it.

Probability scores and uncertainty

Log loss penalizes a wrong prediction more severely as confidence increases. Brier score is the average squared error of the predicted probability. Lower values are better for both. Adding argument to the background vote model improves log loss by 0.0269, with a paired 95% interval from 0.0151 to 0.0379. The interval remains positive when entire terms are resampled.

Argument alone also improves log loss over the historical constant vote rate:0.0317, interval 0.0216 to 0.0415. For the combined model against that constant, the log-loss interval crosses zero. The combined model’s accuracy and Brier improvements over the constant remain positive. These distinctions prevent the weak background model from being the sole benchmark.

Same-sample vote predictions, OT2007 to OT2024
ModelVotesAccuracyLog lossBrier
constant887862.4%0.66230.2347
background887861.2%0.67790.2387
argument887865.0%0.63060.2192
combined887864.9%0.65090.2241

The primary intervals use 4,000 paired bootstrap draws, seed 20261002. Each draw resamples whole SCDB decisions and retains all their dockets and justice votes. Point estimates come from the original sample. The time sensitivity resamples 18 entire terms. It preserves dependence inside each term, but doesn’t model longer serial dependence across terms.

These intervals condition on the fitted historical predictions. Models aren’t refitted inside the bootstrap, so the intervals don’t represent every source of model-development uncertainty. The analysis specifies its primary comparisons and reports the sensitivities together. Per-justice and subset estimates are exploratory and aren’t a new search for significant winners.

Two ways to predict the Court

The aggregated model computes the probability of a strict majority of the scored justices supporting the petitioner, assuming their votes are independent conditional on the measured features. It uses the available votes, which often exclude a silent justice. It isn’t a full-bench forecast. In even-sized sets a tie isn’t a petitioner majority. The accounting in Exhibit 4 compares the background and combined models on this same procedure.

The direct case models use the source court, jurisdiction, United States indicators, and earlier case-win rates, plus the means and standard deviations of the scored justices’ background features. The version with argument adds the means of the seven argument measurements and the standard deviation and median of individual word imbalance. It is trained directly on the case winner. No predicted vote probabilities enter this direct model, and no realized vote margins enter either model.

Direct case prediction changes the estimation target as well as the aggregation method. Its result is a robustness check, not an experiment identifying whether dependence among justices caused the earlier failure. Adding argument yields65.8% accuracy against 66.5% for the direct background model. The paired 95% interval for its accuracy advantage is -2.6 to 0.9 percentage points.

The direct model’s log-loss improvement over its background counterpart is0.0051, interval -0.0117 to 0.0215. Against a probability set to the petitioner win rate in earlier training cases, its log loss is worse. The deterministic “always petitioner” rule is compared on accuracy only; it isn’t treated as a claim of 100% probability.

Stability, coverage, and source quality

Effect of adding argument, with paired 95% log-loss intervals
TestVotesVote accuracy gainVote log-loss intervalCasesCase accuracy gainCase log-loss interval
OT2007–201545784.2 points0.0008 to 0.0402606-1.7 points-0.0355 to 0.0208
OT2016–202443003.3 points0.0217 to 0.04475300.2 points0.0022 to 0.0350
penalty-188783.7 points0.0142 to 0.03781136-0.7 points-0.0229 to 0.0200
penalty-10088783.0 points0.0203 to 0.03941136-1.3 points-0.0003 to 0.0197
penalty-10-retrospective-issue88783.9 points0.0153 to 0.03871136-2.0 points-0.0120 to 0.0221

The vote-level gain persists in both periods and both penalty sensitivities. In OT2016 to OT2024, the direct case model also improves the probability score over its background counterpart. That period doesn’t establish a winner-accuracy advantage. The full-period case probability comparison remains inconclusive.

The issue-area sensitivity adds SCDB’s issue classification to both models. The codebook identifies issues using the Court’s own account of the decision. It is a retrospective adjustment, not verified pre-argument information, and is excluded from the primary test.

Diagnostic subsets; each comparison uses its own matched sample
SubsetCasesVote accuracy gainDirect case accuracy gainCase accuracy interval
Complete recorded bench2472.5 points-0.4 points-2.4 to 1.2 percentage points
Incomplete recorded bench8894.2 points-0.9 points-3.1 to 1.2 percentage points
Unanimous decisions481-3.3 points0.8 points-2.1 to 3.7 percentage points
Divided decisions6558.8 points-2.0 points-4.3 to 0.3 percentage points
One-vote margins20812.4 points-3.8 points-8.2 to 0.0 percentage points
Both transcript sides resolved9785.1 points1.1 points-0.4 to 2.8 percentage points

Full recorded coverage means the number of distinct scored justices equals the reported majority plus minority count. It holds in 247 test cases. The broader sample leaves out votes without side-attributed argument measurements; missing measurements aren’t imputed as silence. The transcript sensitivity requires positive word totals for both sides and no section with an unresolved side. It checks saved predictions on that subset rather than retraining the models. Unanimity and close margins are known only after decision and serve as diagnostics, never prediction inputs.

The input audit found one Oyez argument joined to two separate Medellín decisions. The new test excludes that ambiguous argument. Identical consolidated-docket records are deduplicated to one row per justice and argument; conflicting records would exclude the argument. The input audit collapsed 8 duplicate rows and found no remaining conflicting vote records. Related dockets share a decision-level resampling cluster.

Oyez-derived and SCDB-derived labels agreed on 96.0% of 9,771 overlapping historical votes in the earlier source comparison. Disagreement is a quality concern, not a ceiling on accuracy. The new vote models use SCDB outcomes throughout. The word-counting comparison retains its separately documented source order. A full independent review of every outcome hasn’t been completed.

Reproducing the revision

The research design was written before fitting the new models and amended after the source audit to remove lower-court direction. The read-only export produces a compressed, hashed snapshot of public inputs. The analysis runs locally from that snapshot using Python, NumPy, and SciPy; it needs no live database connection.

The earlier test contained 9,385 votes in 1,201 cases, OT2006 to OT2024. Its 119 changed case calls, 55 corrections, and 64 spoiled calls are reproducible under that specification. They aren’t the revised model’s results. The new sample begins one term later and removes ambiguous and duplicate joins. The historical results remain archived rather than being silently overwritten.

Changes in membership matter to the party-label trend

The descriptive measure asks whether a Republican appointee cast a vote coded conservative, or a Democratic appointee cast one coded liberal. Removing Stevens and Souter largely eliminates the increase across this period. That demonstrates sensitivity to membership. It doesn’t establish that no continuing justice changed, or that the Court’s broader polarization stayed constant.

Supplement 1. Party-label accuracy is sensitive to Court membership

50%68%85%OT05OT15OT24

Every justiceExcluding Stevens and SouterOT2005 to OT2024

Each point is one October Term, divided cases only. The solid line is every justice on the bench. The dashed line removes John Paul Stevens and David Souter, neither of whom sat after OT2009.
Read this as a table
Party-label accuracy in divided cases, by term
TermVotesLabel correctExcluding Stevens and SouterVotes
OT200538468.0%80.5%297
OT200644564.0%76.2%345
OT200747960.5%72.2%371
OT200853965.7%77.6%419
OT200945159.6%64.3%400
OT201041369.5%69.5%413
OT201144770.2%70.2%447
OT201238168.5%68.5%381
OT201331162.4%62.4%311
OT201444070.2%70.2%440
OT201536562.2%62.2%365
OT201625974.9%74.9%259
OT201739476.6%76.6%394
OT201838874.2%74.2%388
OT201943167.3%67.3%431
OT202034666.8%66.8%346
OT202144779.2%79.2%447
OT202232671.2%71.2%326
OT202335974.7%74.7%359
OT202435172.1%72.1%351

The individual rates explain why those two names matter. Their appointing party was usually a poor guide to their ideological votes during the years in this sample. Historical voting records are a more informative baseline for them.

Supplement 2. Stevens and Souter usually voted against their appointment labels

Share of divided-case votes correctly described by appointing party, OT2005 to OT2024. Lines are 95% Wilson intervals. Justices with fewer than 100 votes appear only in the table.
Read this as a table
Vote direction and label accuracy by justice, October Terms 2005 to 2024
JusticeAppointed byTermsDivided votesLabel correct
StevensGerald FordOT2005 to OT200925819.4%
SouterGeorge H. W. BushOT2005 to OT200820827.4%
KennedyRonald ReaganOT2005 to OT201760259.5%
RobertsGeorge W. BushOT2005 to OT202489363.2%
GorsuchDonald J. TrumpOT2016 to OT202434364.4%
BarrettDonald J. TrumpOT2020 to OT202420165.7%
KavanaughDonald J. TrumpOT2018 to OT202429466.3%
BreyerBill ClintonOT2005 to OT202177668.3%
ScaliaRonald ReaganOT2005 to OT201549371.0%
O'ConnorRonald ReaganOT2005 to OT2005771.4%
KaganBarack ObamaOT2009 to OT202461974.0%
GinsburgBill ClintonOT2005 to OT201969675.7%
AlitoGeorge W. BushOT2005 to OT202487377.8%
SotomayorBarack ObamaOT2009 to OT202468378.5%
ThomasGeorge H. W. BushOT2005 to OT202489679.0%
JacksonJoe BidenOT2022 to OT202411481.6%

The usefulness of an appointment label varies by issue

The same descriptive rule has different accuracy across SCDB issue areas. This comparison covers divided cases since October Term 2015. It describes these samples; it doesn’t establish that the gap would persist on future cases.

Supplement 3. Party-label accuracy varies across issue areas

Divided cases since OT2015, with 95% Wilson intervals. Only issue areas with at least 100 coded votes are shown.
Read this as a table
Party-label accuracy by issue area, divided cases since October Term 2015
Issue areaVotesCasesLabel correct95% interval
Civil Rights71916581.8%78.8% to 84.4%
Unions1332377.4%69.6% to 83.7%
Due Process1493177.2%69.8% to 83.2%
Criminal Procedure90725273.6%70.7% to 76.4%
Economic Activity78219167.4%64.0% to 70.6%
Judicial Power4069565.0%60.3% to 69.5%
First Amendment2496261.4%55.3% to 67.3%

Questioning similarity and voting agreement became less closely related

The relationship between how similarly justice pairs questioned and how often they voted together weakened across the measured eras. A comparison restricted to the 6 justices present in both periods also weakens. This is a descriptive association between pairwise measures. It isn’t a direct test of whether the vote-prediction model became less accurate over time.

Supplement 4. Questioning similarity and voting agreement have a weaker relationship in later years

Loading the corpus.
Each dot is a pair of justices over a trailing two-year window. The axes show questioning similarity and agreement in divided votes. The time series tracks their correlation: 0.89 in OT2006-2013, and 0.72 in OT2018-2025. These are correlations, not prediction accuracies.

The move to telephone arguments in 2020 and the later use of ordered questioning could affect these measures, but the analysis doesn’t isolate a cause. The validation ledger reports the alternative samples. The argument ledger also contains the separate seniority-order counts by era and term.

Correction: missing outcomes were presented as undecided cases

An earlier version of the essay presented 22 model outputs as forecasts published before the answers. That claim is withdrawn. The saved file was generated on July 1, 2026, and selected cases without a resolved outcome in its source. That condition doesn’t establish that a case was still undecided.

Two examples establish the problem. Ellingburg v. United States was decided on January 20, 2026. Urias-Orellana v. Bondi was decided on March 4, 2026. Both appear in the July file as awaiting a public outcome. These are verified examples, not an exhaustive audit of the file.

The original file is preserved as a record of the exploratory model outputs, including its incorrect status fields. It shouldn’t be used as evidence of performance on predictions made before decisions. The displayed forecast table and its early-publication claims have been removed. The historical tests described above are separate evaluations.

A future prospective test needs verified decision status, complete accounting for participating justices, a fixed protocol, and a timestamped prediction published before the decision. It must then be scored against independently checked outcomes. No such prospective score is claimed here.

Read the source turns

This explorer holds 115 turns from the two arguments quoted in the essay, covering three justices. It lets you inspect the text behind the examples. Its phrase-rule categories are provisional navigation aids and aren’t inputs to the predictive models.

A blinded second pass by another AI model agreed with the provisional categories on 1 complete label pair out of 30 sampled turns. No human validation is recorded. Neither model’s categories can support a claim about a justice’s legal reasoning on that evidence.

Supplement 5. Each mark opens a source passage

115 of 115 turns shown

Marks sit inside the selected justices’ final vote cells.

Sotomayor · Barack ObamaGorsuch · Donald J. TrumpBarrett · Donald J. Trump
Filter cases and justices
Cases
Justices
Read the accessible legal-grammar table

These counts are modeled by phrase rules. Each count links to supporting turn IDs in the versioned exhibit snapshot.

Provisional legal-grammar counts for the selected source slice
JusticeEvidence objectLegal operationTurns
Sotomayorcase record or adjudicative factstest a factual assumption1
Sotomayorinstitutional authority or competencechallenge a premise1
Sotomayorprocedure or jurisdictionclarify a premise2
Sotomayorprocedure or jurisdictionidentify a procedural gate4
Sotomayorprocedure or jurisdictionnarrow or broaden a proposed rule1
Sotomayorprocedure or jurisdictionunclassified operation1
Sotomayorquantitative or empirical claimchallenge a premise1
Sotomayorquantitative or empirical claimunclassified operation1
Sotomayorunclassified evidencechallenge a premise5
Sotomayorunclassified evidencedistinguish an authority1
Sotomayorunclassified evidencetest a boundary through a hypothetical1
Sotomayorunclassified evidenceunclassified operation27
Gorsuchcase record or adjudicative factstest a factual assumption3
Gorsuchcase record or adjudicative factsunclassified operation2
Gorsuchconstitutional or statutory textunclassified operation1
Gorsuchinstitutional authority or competenceunclassified operation1
Gorsuchprocedure or jurisdictionchallenge a premise2
Gorsuchprocedure or jurisdictionidentify a procedural gate1
Gorsuchprocedure or jurisdictiontest a factual assumption1
Gorsuchprocedure or jurisdictionunclassified operation3
Gorsuchunclassified evidencechallenge a premise5
Gorsuchunclassified evidenceunclassified operation30
Barrettcase record or adjudicative factstest a factual assumption1
Barrettinstitutional authority or competenceclarify a premise1
Barrettprocedure or jurisdictionidentify a procedural gate1
Barrettprocedure or jurisdictionunclassified operation1
Barrettunclassified evidencechallenge a premise4
Barrettunclassified evidenceclarify a premise2
Barrettunclassified evidencenarrow or broaden a proposed rule1
Barrettunclassified evidencetest a boundary through a hypothetical3
Barrettunclassified evidenceunclassified operation6
Turns from Bowe v. United States and Cox Communications, with source links, observed timing, and provisional modeled categories.

Research and reproducibility