Published standard · Evidence v1
Methodology and limitations
CIP observes sampled answers from specified models, questions, settings, and times. It does not inspect hidden model state or treat one answer as fact.
Evidence grades
Fewer than 20 answers is insufficient evidence. Above that threshold, single-model directional evidence remains distinct from multi-model corroboration. Scores show sample size, intervals, and method version.
Facts and sources
A reachable link proves only that a source exists. Source class and machine-assessed support are not independent certification of truth.
Quasi-experiments
Treatment/control results use question-level aggregation, seeded stratified assignment, and difference estimates. Underpowered or non-preregistered work remains directional.
Known limits
Outputs vary with model version, geography, time, prompts, and provider policy. Every report is a time-bounded sample.