Skard assertion codebook
Version: skard-assertion-codebook-v1
Status: active pilot codebook, 14 September 2026
This codebook governs coverage_assertion records. Its unit is one bounded source statement or inspected page set linked to one atlas concept. Every axis answers a different question. Coders must not compensate for a missing value on one axis by changing another.
The machine-readable vocabularies are in data/topic_graph.json; the JSON Schema requires every assertion to use them. A codebook change that alters a category's meaning creates a new version and requires affected assertions to be reviewed again.
1. Mention status
| Value | Use when | Do not use when |
|---|---|---|
exact_explicit | The bounded source literally names the target or an accepted alias. | The target is only inferred from an actor, date or broader period. |
indirect_explicit | The source explicitly describes the target without using its accepted name. | The connection depends mainly on the atlas graph or outside knowledge. |
not_explicit_in_inspected_scope | A documented, complete page set was checked with recorded queries and no explicit target was found. | Pages are missing, OCR is empty, or the span was not assessed. |
not_assessed | No defensible search/review has been completed. | A zero-result search over a verified span exists. |
Positive example: gnw-1922-explicit literally says “Den store nordiske krigen.” Negative example: gnw-1939-absence records a bounded null finding; it does not claim that teachers omitted the war.
For blind coding, the exported target_aliases field is the complete accepted alias list for that batch. A descriptive phrase absent from both the preferred label and this list is indirect_explicit, even when it maps cleanly to the target definition.
2. Semantic relation
| Value | Meaning |
|---|---|
exact_concept | The source statement denotes the assertion target. |
narrower_instance | The statement denotes one instance or subtopic of the target. |
broader_frame | The source gives a category that could include the target but does not identify it. |
overlapping_frame | Source and target share material without either containing the other. |
analyst_inference | The connection is an explicit editorial interpretation and must not be presented as source wording. |
gnw-2026-primary-absence is a broader frame: “historiske konflikter” can include the Great Northern War, but is not another name for it. Graph relations never populate this field automatically.
3. Salience
| Value | Meaning |
|---|---|
main_focus | The target is the main point of the bounded statement. |
sub_focus | It is a meaningful subordinate point. |
passing_mention | It is named but not developed in that statement. |
opportunity_only | The wording merely creates room in which the target could be selected. |
Salience concerns the bounded evidence, not the whole document and not the historical importance of the topic.
4. Normative force
| Value | Meaning |
|---|---|
required | The relevant content/action is prescribed. |
recommended | The source uses advisory force such as “bør.” |
permitted | It explicitly allows something without recommending it. |
prohibited | It explicitly disallows something. |
unspecified | The bounded evidence does not establish force. |
Normative force never says whether the item is named or freely selected. That belongs to selection mode.
The governing heading or modal may sit outside the short publishable quotation. Use it only when the exported page_context unambiguously governs the quoted statement. If the context does not establish that link, use unspecified.
5. Selection mode
| Value | Meaning |
|---|---|
named_item | A particular event, person, work or other item is prescribed or presented. |
category_open_selection | A category is fixed while concrete examples remain open. |
choice_set | Selection must be made from a bounded set. |
example | The source offers the item as an example. |
not_applicable | Selection is not meaningful for this assertion. |
unclear | The evidence does not support a firmer classification. |
The two axes combine freely. gnw-1922-explicit is required + named_item. The LK20 Romanticism assertions are required + category_open_selection. selection-1974 is recommended + category_open_selection. The separate selection-1974-author-representation assertion shows why a required or recommended category is not a named canon: “the best-known Norwegian literary authors” may be represented while the identity inventory remains zero.
Quantitative requirements
Use optional quantitative_requirement only when the evidence gives a numeric threshold, exact quantity, approximation or distribution. This is a structured facet of an assertion, not an eighth categorical axis. Preserve the source's modal under normative_force; a recommendation of at least 800 pages is not silently promoted to a requirement. Record metric, comparator, value, unit and grade span, plus any explicit distribution such as approximately one third in the secondary written standard. Do not infer a quantity from words such as “many”, “representative” or “substantial”.
6. Target kind
Use one principal kind for the target of this assertion:
substantive_content: events, people, works, periods and other subject matter;disciplinary_knowledge: source criticism, causation, perspective and other subject-specific ways of knowing;generic_competence: broad transferable capacity;method_activity: a prescribed or recommended teaching/learning activity;value_attitude: a value, disposition or attitude;assessment: what or how performance is assessed.
If one sentence supports genuinely different targets, create separate assertions against separate controlled concepts rather than hiding them in one multi-purpose record.
7. Student actions
Record only actions explicit in the bounded evidence, using the controlled verb list. Multiple actions may be attached to one assertion. Preserve the original verb in the quotation; the controlled value is a functional mapping.
Allowed values: identify, recall, narrate, explain, compare, interpret, source_critique, evaluate, explore, reflect, read, create, discuss, participate.
An empty array means “no student action explicitly coded in this evidence.” It does not mean that the wider curriculum contains no student activity.
Examples:
- “utforske og reflektere” →
explore,reflect; - “sammenligne og tolke” →
compare,interpret; - “gjøre rede for” →
explain.
8. Independent coding and agreement
The active primary coding is not an acceptable blind worksheet. Use scripts/export_blind_coding.py to export source context, target definitions, quoted evidence and negative-search scope while withholding every existing axis value and editorial claim. A second coder records their identity and codes the batch without consulting the public assertion cards.
scripts/score_coding.py reports observed agreement and Cohen's kappa for each single-valued axis, linear-weighted kappa for salience, and exact-set agreement plus mean Jaccard similarity for the multi-valued student-action axis. Preserve the completed coding sheet and report before reconciliation. A score does not itself update review state; accepted reconciliations must be new append-only review decisions.
9. Population and comparability
Built assertions carry jurisdiction, authority, administrative level, document layer, school type, programme/track, grade interval, typical-age status, compulsory status, language form and ISCED-mapping status. The historical display label remains alongside these normalized fields.
Unknown values are recorded as unknown, never inferred from the next curriculum or contemporary school structure. An assertion-level override may narrow a document spanning several stages: the Vg2 and Vg3 Romanticism assertions are therefore distinct populations even though both originate in NOR01-08.
Population equality is necessary but not sufficient for a controlled before/ after comparison. Subject layer, authority, validity and source completeness must still be inspected.
10. Review and machine assistance
Search, OCR, embeddings and models may propose records. Only an explicit review decision can make one public as reviewed. A decision pins the codebook-sensitive record revision, source version and evidence anchor. Recoding appends a superseding decision; earlier decisions remain auditable.
single_coded currently means one model-assisted editorial coding. It is not human validation. Before aggregate claims, a stratified sample must be independently coded and reliability reported per axis.
11. Borderline rules
- A named person near an event is a person mention, not automatically an event assertion.
- A required period with open text selection is not a named author or work canon.
- A broad category plus a null search may establish opportunity and non-explicit status simultaneously.
- A title in a table of contents establishes listing, not depth, use or learning.
- Absence, prohibition and
not_assessedare three different findings. - A required category plus an example list produces two facts: the category may be required while each named example remains merely permitted.
- A conditional exam option preserves both levels of choice (
condition_group_idand any inneralternative_group_id); it is not flattened into a universal requirement. - A textbook biography that names a work proves a mention in that textbook, not that the work itself is reproduced, assigned or read.
- A curriculum bibliography or library list proves resource provision within its governing wording. It does not by itself turn the named authors or titles into literary selections, required reading or evidence of classroom use.
- A person proposed as the topic of a reader or booklet is a
resource-topic, not a bibliography creator, historical syllabus requirement or literary author selection. Preserve the surrounding resource recommendation separately. - A catalogue-attested publication/activity window is not a person's lifespan, and a catalogue-attested series window is not a verified first/last publication interval. Preserve the status and source note wherever dates are shown.
- An ambiguous surname or short form remains an unresolved source identity until evidence distinguishes candidates. Do not borrow a famous person's authority identifiers merely because one candidate seems likely.
- Entity zeroes may be displayed only for an
assessedinventory whose scope basis, source anchor, population boundary, completeness decision and every bound member review are current.page_rangemeans the whole declared span;evidence_groupsmeans only the named list contexts and must enumerate them. The two forms cannot be mixed.partialis always unknown. - A work title that is also an ordinary word or an unresolved short form must not drive cross-corpus literal matches. Use qualified
literal_search_aliasesor leave literal matching disabled until contextual review exists.