SKARD English

Skard assertion codebook

Version: skard-assertion-codebook-v1

Status: active pilot codebook, 14 September 2026

This codebook governs coverage_assertion records. Its unit is one bounded source statement or inspected page set linked to one atlas concept. Every axis answers a different question. Coders must not compensate for a missing value on one axis by changing another.

The machine-readable vocabularies are in data/topic_graph.json; the JSON Schema requires every assertion to use them. A codebook change that alters a category's meaning creates a new version and requires affected assertions to be reviewed again.

1. Mention status

ValueUse whenDo not use when
exact_explicitThe bounded source literally names the target or an accepted alias.The target is only inferred from an actor, date or broader period.
indirect_explicitThe source explicitly describes the target without using its accepted name.The connection depends mainly on the atlas graph or outside knowledge.
not_explicit_in_inspected_scopeA documented, complete page set was checked with recorded queries and no explicit target was found.Pages are missing, OCR is empty, or the span was not assessed.
not_assessedNo defensible search/review has been completed.A zero-result search over a verified span exists.

Positive example: gnw-1922-explicit literally says “Den store nordiske krigen.” Negative example: gnw-1939-absence records a bounded null finding; it does not claim that teachers omitted the war.

For blind coding, the exported target_aliases field is the complete accepted alias list for that batch. A descriptive phrase absent from both the preferred label and this list is indirect_explicit, even when it maps cleanly to the target definition.

2. Semantic relation

ValueMeaning
exact_conceptThe source statement denotes the assertion target.
narrower_instanceThe statement denotes one instance or subtopic of the target.
broader_frameThe source gives a category that could include the target but does not identify it.
overlapping_frameSource and target share material without either containing the other.
analyst_inferenceThe connection is an explicit editorial interpretation and must not be presented as source wording.

gnw-2026-primary-absence is a broader frame: “historiske konflikter” can include the Great Northern War, but is not another name for it. Graph relations never populate this field automatically.

3. Salience

ValueMeaning
main_focusThe target is the main point of the bounded statement.
sub_focusIt is a meaningful subordinate point.
passing_mentionIt is named but not developed in that statement.
opportunity_onlyThe wording merely creates room in which the target could be selected.

Salience concerns the bounded evidence, not the whole document and not the historical importance of the topic.

4. Normative force

ValueMeaning
requiredThe relevant content/action is prescribed.
recommendedThe source uses advisory force such as “bør.”
permittedIt explicitly allows something without recommending it.
prohibitedIt explicitly disallows something.
unspecifiedThe bounded evidence does not establish force.

Normative force never says whether the item is named or freely selected. That belongs to selection mode.

The governing heading or modal may sit outside the short publishable quotation. Use it only when the exported page_context unambiguously governs the quoted statement. If the context does not establish that link, use unspecified.

5. Selection mode

ValueMeaning
named_itemA particular event, person, work or other item is prescribed or presented.
category_open_selectionA category is fixed while concrete examples remain open.
choice_setSelection must be made from a bounded set.
exampleThe source offers the item as an example.
not_applicableSelection is not meaningful for this assertion.
unclearThe evidence does not support a firmer classification.

The two axes combine freely. gnw-1922-explicit is required + named_item. The LK20 Romanticism assertions are required + category_open_selection. selection-1974 is recommended + category_open_selection. The separate selection-1974-author-representation assertion shows why a required or recommended category is not a named canon: “the best-known Norwegian literary authors” may be represented while the identity inventory remains zero.

Quantitative requirements

Use optional quantitative_requirement only when the evidence gives a numeric threshold, exact quantity, approximation or distribution. This is a structured facet of an assertion, not an eighth categorical axis. Preserve the source's modal under normative_force; a recommendation of at least 800 pages is not silently promoted to a requirement. Record metric, comparator, value, unit and grade span, plus any explicit distribution such as approximately one third in the secondary written standard. Do not infer a quantity from words such as “many”, “representative” or “substantial”.

6. Target kind

Use one principal kind for the target of this assertion:

  • substantive_content: events, people, works, periods and other subject matter;
  • disciplinary_knowledge: source criticism, causation, perspective and other subject-specific ways of knowing;
  • generic_competence: broad transferable capacity;
  • method_activity: a prescribed or recommended teaching/learning activity;
  • value_attitude: a value, disposition or attitude;
  • assessment: what or how performance is assessed.

If one sentence supports genuinely different targets, create separate assertions against separate controlled concepts rather than hiding them in one multi-purpose record.

7. Student actions

Record only actions explicit in the bounded evidence, using the controlled verb list. Multiple actions may be attached to one assertion. Preserve the original verb in the quotation; the controlled value is a functional mapping.

Allowed values: identify, recall, narrate, explain, compare, interpret, source_critique, evaluate, explore, reflect, read, create, discuss, participate.

An empty array means “no student action explicitly coded in this evidence.” It does not mean that the wider curriculum contains no student activity.

Examples:

  • “utforske og reflektere” → explore, reflect;
  • “sammenligne og tolke” → compare, interpret;
  • “gjøre rede for” → explain.

8. Independent coding and agreement

The active primary coding is not an acceptable blind worksheet. Use scripts/export_blind_coding.py to export source context, target definitions, quoted evidence and negative-search scope while withholding every existing axis value and editorial claim. A second coder records their identity and codes the batch without consulting the public assertion cards.

scripts/score_coding.py reports observed agreement and Cohen's kappa for each single-valued axis, linear-weighted kappa for salience, and exact-set agreement plus mean Jaccard similarity for the multi-valued student-action axis. Preserve the completed coding sheet and report before reconciliation. A score does not itself update review state; accepted reconciliations must be new append-only review decisions.

9. Population and comparability

Built assertions carry jurisdiction, authority, administrative level, document layer, school type, programme/track, grade interval, typical-age status, compulsory status, language form and ISCED-mapping status. The historical display label remains alongside these normalized fields.

Unknown values are recorded as unknown, never inferred from the next curriculum or contemporary school structure. An assertion-level override may narrow a document spanning several stages: the Vg2 and Vg3 Romanticism assertions are therefore distinct populations even though both originate in NOR01-08.

Population equality is necessary but not sufficient for a controlled before/ after comparison. Subject layer, authority, validity and source completeness must still be inspected.

10. Review and machine assistance

Search, OCR, embeddings and models may propose records. Only an explicit review decision can make one public as reviewed. A decision pins the codebook-sensitive record revision, source version and evidence anchor. Recoding appends a superseding decision; earlier decisions remain auditable.

single_coded currently means one model-assisted editorial coding. It is not human validation. Before aggregate claims, a stratified sample must be independently coded and reliability reported per axis.

11. Borderline rules

  • A named person near an event is a person mention, not automatically an event assertion.
  • A required period with open text selection is not a named author or work canon.
  • A broad category plus a null search may establish opportunity and non-explicit status simultaneously.
  • A title in a table of contents establishes listing, not depth, use or learning.
  • Absence, prohibition and not_assessed are three different findings.
  • A required category plus an example list produces two facts: the category may be required while each named example remains merely permitted.
  • A conditional exam option preserves both levels of choice (condition_group_id and any inner alternative_group_id); it is not flattened into a universal requirement.
  • A textbook biography that names a work proves a mention in that textbook, not that the work itself is reproduced, assigned or read.
  • A curriculum bibliography or library list proves resource provision within its governing wording. It does not by itself turn the named authors or titles into literary selections, required reading or evidence of classroom use.
  • A person proposed as the topic of a reader or booklet is a resource-topic, not a bibliography creator, historical syllabus requirement or literary author selection. Preserve the surrounding resource recommendation separately.
  • A catalogue-attested publication/activity window is not a person's lifespan, and a catalogue-attested series window is not a verified first/last publication interval. Preserve the status and source note wherever dates are shown.
  • An ambiguous surname or short form remains an unresolved source identity until evidence distinguishes candidates. Do not borrow a famous person's authority identifiers merely because one candidate seems likely.
  • Entity zeroes may be displayed only for an assessed inventory whose scope basis, source anchor, population boundary, completeness decision and every bound member review are current. page_range means the whole declared span; evidence_groups means only the named list contexts and must enumerate them. The two forms cannot be mixed. partial is always unknown.
  • A work title that is also an ordinary word or an unresolved short form must not drive cross-corpus literal matches. Use qualified literal_search_aliases or leave literal matching disabled until contextual review exists.