N1 2017-12 — audit report
70 questions (問題1–13, the non-listening half). Generated from the OCR and correction audit trails.
OCR confidence
Source: retypeset-digital — quality 3/5: Clean, fully legible retypeset digital text (PDF has a complete embedded text layer matching the page images), with underlines for target words preserved. Marked down for: bailitop Chinese ad header/footer on every page plus a large gray watermark crossing the text; loss of original layout (no boxed question numbers, re-flowed line breaks, options split across pages); two content elements surviving only as lower-quality embedded scans (問題13 tables on p24, slightly cropped, and the 聴解問題1 2番 map on p26); scattered retypesetting typos (バスワード, 再立付, 所届学部, ロボット論理, 選らんで, 確認者 for 確認書 in Q46 option 1); and the missing listening 問題5 3番 printed options.
Every section was transcribed twice independently (a full-page pass and a double-resolution half-page pass) and reconciled. Of 13 sections, 2 agreed on all content between the two passes (reconciled automatically, differing only in transcriber notes) and 11 had at least one content difference resolved by a third adjudication pass that re-read the page images. 38 leaf-level differences were examined in total.
Residual OCR risks (19):
- 問題1: transcribed from text witnesses, not the booklet image; see the section-specific audit notes
- 問題2: transcribed from text witnesses, not the booklet image; see the section-specific audit notes
- 問題3: transcribed from text witnesses, not the booklet image; see the section-specific audit notes
- 問題4: transcribed from text witnesses, not the booklet image; see the section-specific audit notes
- 問題5: transcribed from text witnesses, not the booklet image; see the section-specific audit notes
- 問題6: transcribed from text witnesses, not the booklet image; see the section-specific audit notes
- 問題7: transcribed from text witnesses, not the booklet image; see the section-specific audit notes
- 問題8: The こ-for-ご typos (こざいます/こ希望) and the redundant 宛あて are preserved as printed; they are genuine source-retype artifacts, not OCR errors, confirmed against the high-res half-page renders.
- 問題8: Passage labels (一)/(2)/(3)/(4) are inconsistent in the source but dropped per output shape; only noted.
- 問題9: The exact presence/absence of a single space after each (注N) marker is hard to read under the watermark; followed the clearest reading (space after passage 2 注1 and passage 3 注1, none elsewhere). Low impact.
- 問題9: The long dash in 育んだ産業― rendered as a single U+2015 as printed (it appears as one dash, not the ―― 2-em form); kept as both transcribers had it.
- 問題9: The half-width '?' after でしょうか? / のでしょうか? is kept exactly as printed (the source uses ASCII ?, not full-width ?); both transcribers agreed.
- 問題10: no human/agent spot-check was performed for this section (A/B byte-agreement on content was treated as sufficient)
- 問題11: The opening and closing quote glyphs around おれもやれるな are both rendered as slanted double-prime marks; transcribed as U+2033 on both sides. B noted a left/right directional difference, but at available resolution both read as U+2033; the precise intended codepoint (″ vs “”) is a retypeset artifact.
- 問題11: The half-width vs full-width parentheses distinction for (注) inline vs (注)gloss follows the print as rendered; OCR cannot fully rule out a font-metric rendering artifact, but the inline mark reads consistently narrower.
- 問題12: no human/agent spot-check was performed for this section (A/B byte-agreement on content was treated as sufficient)
- 問題13: Table-2 5-a middle character: the raster best-reads as 調育書 (月 component visible) but the heavy scan leaves residual doubt between 育 and a degraded 査; recorded as 調育書 with a note flagging the 調査書/調書 inconsistency across the document.
- 問題13: Table cell spanning is modeled with plain <td>s (content in the first row of a merged group, blanks below); the exact vertical extent of the merged 受付時間/所要日数 cells is inferred from divider lines in a low-resolution raster.
- 問題13: メ-ル/センタ- are transcribed with a half-width hyphen as literally printed; if the project later prefers normalizing the obvious intended long-vowel mark ー, these would change.
Source-text quirks the OCR transcribed verbatim (printed typos/recall artifacts; the ones judged to be retype errors were corrected in the layer below) — 12 noted at OCR time:
- 問題8: Passage (1) email contains likely retype typos, transcribed as printed: 「ありがとうこざいます」 (こ for ご in ございます), 「こ希望の便」「こ希望に添えなかった」 (こ for ご in ご希望), and redundant 「ご自宅宛あてに」 (宛 immediately followed by あて).
- 問題8: Passage (2) and its (注) gloss both spell the reading of 渇望 as かつぼう; transcribed as printed.
- 問題9: The retype prints inline kana readings as HALF-WIDTH parentheses on the baseline (e.g. 物見(ものみ)遊山(ゆさん), 富士山(ふじさん), 稀(まれ), 天邪鬼(あまのじゃく), 包摂(ほうせつ), 依拠(いきょ)); normalized to
<ruby>markup in output to match JLPT ruby style. The full-width section markers (1)(2)(3) are omitted as passage labels. - 問題11: The (注)こなす:処理する gloss is printed below passage B but glosses the word こなせる in passage A; placed at end of passage B as printed.
- 問題11: Passage A's こなせる(注) uses a half-width parenthesized 注 printed inline; transcribed as printed.
- 問題12: The 注 gloss line is printed as a partial re-typeset that erroneously repeats the inline furigana and the (注) marker: it reads '(注)空疎(くうそ)(注)な常套句(じょうとうく):ここでは、中身のない形だけの言葉'. Normalized in output to ruby and with the duplicated inline (注) removed.
- 問題12: Possible source typo: 'ロボット論理' (リ) where 'ロボット倫理' might be expected given the surrounding list of ○○倫理 terms; transcribed as printed.
- 問題12: The 注 gloss as printed has a duplicated 注 marker and a doubled occurrence of 空疎: it reads 「(注)空疎(くうそ)(注)な常套句(じょうとうく):…」. This is a retype artifact (the original passage has 空疎(くうそ)(注)な常套句(じょうとうく), where 常套句 is the term being glossed, not 空疎), so the output normalizes the gloss to 「(注)空疎な常套句:…」 with ruby.
- 問題12: In the passage body, the inline furigana parentheses 空疎(くうそ) and 常套句(じょうとうく) were converted to
<ruby>form. The retype's note marker placement is corrected in the layer below to follow 常套句. - 問題13: Table 2 raster oddities: the row label '5-a' uses a long full-width dash while '5-b' uses a short half-width hyphen; transcribed as seen. The middle character of the 5-a item reads 調育書 in the raster (three characters, middle has a 月 component) — a likely retype garble; 5-b reads 調書 (two characters). Note the question 69 stem (clean digital text) writes 調査書, so the source is internally inconsistent. All kept exactly as printed in each location.
- 問題13: 再立付 (in パスワードを忘れた…再立付を願い出てください) is an unusual/likely-garbled retypeset word (perhaps for 再発行); transcribed as printed.
- 問題13: メ-ル and センタ- in the clean digital body use a short half-width hyphen rather than the long vowel mark ー; transcribed as printed. (Note: the question-70 answer text uses a real long vowel mark, メール, and is kept as such.)
Corrections applied (10)
Evidence-based reconstructions of the original exam text. This section is the canonical correction log for the sitting.
- 問3 q17
question: 「<u>撤回</u>した」 → 「<u>撤回した</u>」- All four choices already end in した, so leaving the source retype's した outside the tested span duplicates the inflection. The full target is 撤回した.
- evidence: structural: 問題3 choices replace the underlined span.
- 問3 q18
question: 「<u>張り合っている</u>」 → 「<u>張り合って</u>いる」- All four choices are て-forms, so いる must remain outside the tested span for grammatical substitution (競い合っている). The source retype extended the underline too far, exactly as in the repeated N1 2010-07 q16 item.
- evidence: structural: 問題3 choices replace the underlined span; repeated item in N1 2010-07 q16.
- 問8 q46
passage: 「ありがとうこざいます」 → 「ありがとうございます」- こざいます is not a Japanese word; the polite verb is ございます. A retype こ-for-ご artifact (quirks confirms it is a retype artifact, which is exactly this task's target). The true exam read ありがとうございます.
- evidence: Category typo, confidence high. [retype-corrections-pass2 (propose + adversarial verify), 2026-06-20]
- 問8 q46
passage: 「こ希望の便が確保」 → 「ご希望の便が確保」- こ希望 is not a word; the honorific prefix is ご (ご希望). Retype こ-for-ご artifact. True exam read ご希望.
- evidence: Category typo, confidence high. [retype-corrections-pass2 (propose + adversarial verify), 2026-06-20]
- 問8 q46
passage: 「万が一、こ希望に添えなかった」 → 「万が一、ご希望に添えなかった」- こ希望 is not a word; honorific prefix ご (ご希望). Retype こ-for-ご artifact. True exam read ご希望.
- evidence: Category typo, confidence high. [retype-corrections-pass2 (propose + adversarial verify), 2026-06-20]
- 問8 q46
passage: 「ご自宅宛あてに確認書」 → 「ご自宅あてに確認書」- 宛あて redundantly writes the same word twice (宛 plus its kana reading あて). Polite business Japanese normally writes ご自宅あてに in kana; the doubled 宛 is a retype artifact.
- evidence: Category added-char, confidence medium. [retype-corrections-pass2 (propose + adversarial verify), 2026-06-20]
- 問12 q65
passage: 「ロボット論理といった」 → 「ロボット倫理といった」- The entire list enumerates ○○倫理 terms (生命倫理、環境倫理、情報倫理、工学倫理…遺伝子倫理、脳神経倫理、ナノ倫理); 論理 (logic, りんり homophone of 倫理) breaks the pattern and must be ロボット倫理.
- evidence: Category homophone, confidence high. [retype-corrections-pass2 (propose + adversarial verify), 2026-06-20]
- 問13 q69
passage: 「5-a.調育書(北山大学の書式)」 → 「5-a.調査書(北山大学の書式)」- 調育書 is not a real document name; the matching 5-b cell is 調書 and question 69 references 調査書(書式は会社指定のもの). 育 is a misread/retype of 査; the intended term is 調査書.
- evidence: Category typo, confidence high. [retype-corrections-pass2 (propose + adversarial verify), 2026-06-20]
- 問13 q69
passage: 「再立付を願い出て」 → 「再発行を願い出て」- 再立付 is not a Japanese word; in context (forgotten password, request a new one) the intended term is 再発行 (reissue). A garbled retype.
- evidence: Category typo, confidence medium. [retype-corrections-pass2 (propose + adversarial verify), 2026-06-20]
- 問12 q65
passage:<ruby>空疎<rt>くうそ</rt></ruby>(注)な<ruby>常套句<rt>じょうとうく</rt></ruby>にすぎない→<ruby>空疎<rt>くうそ</rt></ruby>な<ruby>常套句<rt>じょうとうく</rt></ruby>(注)にすぎない- (注N)note marker repositioned to follow its glossed term (the retype placed the marker before the word).
- evidence: Marker-only move; the marker-stripped text is byte-identical. [note-marker relocation pass, 2026-06-20]
Answer Confidence
69 high · 0 medium · 0 low · 1 unresolved (of 70).
Generated from answer-audit.json at build time. Questions listed below are below high confidence or have source disagreement.
Answer Sources
bailitop(key) — N1_2017-12_kaisetsu.pdfjlpt-freq/extracted_text/N1/2017/2017.12_n1-2017.12真题+答案+听力原文.txt(key) — github.com/5Mcv4He/jlpt-freq extracted_text/N1/2017/2017.12_n1-2017.12真题+答案+听力原文.txtsolver-1(solver)solver-2(solver)
Questions Needing Attention
| Question | Chosen | Confidence | Notes | Votes |
|---|---|---|---|---|
| q36 | unresolved | unresolved | bailitop=1, jlpt-freq/text/N1/2017/2017.12_n1-2017.12真题+答案+听力原文=1, solver-1=4, solver-2=4 |