N1 2023-12 — audit report
66 questions (問題1–13, the non-listening half). Generated from the OCR and correction audit trails.
OCR confidence
Source: booklet-scan — quality 4/5: Clean, high-contrast scan of the original printed booklets; body text, boxed question numbers, and furigana are all crisply legible. Every page carries a light diagonal Vietnamese watermark ('Tôi Yêu Ngoại Ngữ Group / Yuuki Bùi') across the upper-left, overlapping text on most pages without impairing legibility. Minor issues: right-edge section tabs (文法/読解) are clipped on a few pages, faint halftone grain/dust specks throughout, and a small scanner smudge on listening page 2; not a retypeset digital text, so short of a 5.
Every section was transcribed twice independently (a full-page pass and a double-resolution half-page pass) and reconciled. Of 13 sections, 0 agreed on all content between the two passes (reconciled automatically, differing only in transcriber notes) and 13 had at least one content difference resolved by a third adjudication pass that re-read the page images. 60 leaf-level differences were examined in total.
Calibration: mondai1–6 matched the independently hand-curated reference on 196/200 fields; all four differences were errors in the reference, not the OCR.
Residual OCR risks (34):
- 問題1: The conventions' rule text (per-kanji ruby when the reading split is unambiguous) is in tension with its whole-word surname example (池田→いけだ) for question 3's 村上/むらかみ; the whole-word surname exemplar was followed, but a reviewer applying the rule text literally could prefer A's per-kanji split.
- 問題1: The ___ representation of the instruction's blank underline is an uncovered-by-convention judgment call both transcribers happened to share; downstream consumers should treat its exact width as arbitrary.
- 問題2: The stray hyphen-like mark inside the Q13 blank was normalized away per the blank convention; if it were ever deemed intentional printed content the transcription would need revisiting, but at full resolution it reads as a typesetting/scan artifact.
- 問題3: The ___ rendering of the instruction's leading underline rule is an ad-hoc convention-gap decision; other sections of this exam (and other exams) must use the same rendering for cross-section consistency.
- 問題3: Q17's inclusion of 。 inside <u>…</u> rests on reading the underline endpoint at double resolution; the reading is clear, but it is the only place in this section where a one-character-cell judgement changes the transcription.
- 問題3: Q16's whole-word ruby deliberately diverges from the printed per-kanji furigana placement because the conventions hard-code 池田 as the whole-word example; if the convention is later reinterpreted as per-kanji-when-split-is-printed, this field would change.
- 問題4: The Q22 options 1-3 underline stopping mid-conjugation (before いる/いた) is typographically unusual for this question type, but pixel measurement is unambiguous; if a higher-fidelity source ever surfaces, this is the one spot worth re-confirming.
- 問題4: Both transcribers could share a correlated error on lines not spot-checked, but six lines plus all diff sites were verified character-by-character against the double-resolution scans with no discrepancies found.
- 問題5: 山下 furigana is transcribed whole-word (<ruby>山下<rt>やました</rt></ruby>); the printed rubies span the whole name, and a per-kanji split 山(やま)下(した) would also be defensible, but both transcribers used whole-word per the name example (池田) in the conventions.
- 問題5: Q31 option 3 「買ったかにかかわらず」 is unusual Japanese and likely a retypesetting artifact in the source booklet; it is transcribed as printed per the conventions, so downstream consumers should not 'correct' it.
- 問題6: The printed slot blanks in the booklet are underlines of varying width; per conventions they are normalized to the ' _ _ ★ _ ' pattern, so the transcription does not preserve the printed blank widths (intentional).
- 問題6: p11-b.png is blank except the page number, confirming questions 39-40 are the entirety of mondai6's second page; no content was missed there.
- 問題7: The furigana over 鶯 is printed very small; read as うぐいす at 4x zoom and both transcribers agree, but individual glyph quality is marginal.
- 問題7: The lead-in line 「以下は、作家が書いたエッセイである。」 is printed outside the passage box but transcribed as the first <p> of the passage; downstream consumers should be aware of this placement convention.
- 問題8: 活き活き is printed with furigana い over each 活 only (き is okurigana); per the whole-word convention it was transcribed as <ruby>活き活き<rt>いきいき</rt></ruby>, which folds the okurigana into the ruby base. If the convention is later interpreted as 'whole kanji run' rather than 'whole word including okurigana', this should become <ruby>活<rt>い</rt></ruby>き<ruby>活<rt>い</rt></ruby>き.
- 問題8: The gap between 9月13日 and 13:30 in the 日時 cell was judged a full-width space from visual width; a half-width space is conceivable.
- 問題8: The justification spaces printed inside the header labels (件 名, 日 時) are dropped in the <table> labels per B's rendering; if header labels are later compared verbatim against the print, this normalization should be remembered.
- 問題8: correctAnswerIndex is null for all four questions; no answer key appears on these pages.
- 問題8: The diagonal Vietnamese watermark crosses several lines on p14-p17; all affected glyphs were legible at double resolution, but the watermark slightly degrades confidence on small furigana.
- 問題9: The 35〜40 wave dash in passage (2) is transcribed as 〜 (U+301C) per both transcribers; the printed glyph could be a fullwidth tilde ~ (U+FF5E). Both agreed, so it was not re-adjudicated.
- 問題9: Inline (注N) marker glyphs are tiny in the scan; their exact horizontal placement relative to surrounding kana (e.g., 鵜呑み(注1)に vs 鵜呑(注1)み) was judged by best reading and the mondai7/8 precedent, but sub-character placement carries minor uncertainty.
- 問題9: The full-width spaces after ? and ! in passage (3) are inferred from inter-glyph gaps in the scan; rendering ambiguity is possible though both transcribers and the source layout support them.
- 問題10: The printed parentheses around (中略) and the in-text (注) mark are small enough that half-width vs full-width cannot be distinguished with certainty; transcribed full-width per the convention that Japanese punctuation stays full-width, and both transcribers agree.
- 問題10: The leader after 「シュバイッツア博士」 is transcribed as …… (six-dot leader, two ellipsis characters); the exact dot count is consistent with the high-resolution image but small print leaves minor uncertainty.
- 問題10: The colon in the 注 gloss (注)アングル:角度 is transcribed as full-width :; at print size a half-width colon would look nearly identical.
- 問題11: The furigana りちぎ is printed as a continuous run over 律儀; the per-kanji split 律(りち)儀(ぎ) follows the standard reading and seems unambiguous, but the print itself does not mark the boundary.
- 問題11: The diagonal watermark crosses several lines of passage A on p28; all overlapped characters were legible at double resolution, but the watermark slightly lowers contrast around 律儀さ and キリがない.
- 問題11: The phrasing 「…と考える人は「小説の書き方マニュアル」を信じる律儀さと同じで」 reads awkwardly and may itself be a retypeset artifact of the leaked booklet; it is transcribed exactly as printed per the conventions.
- 問題12: The exact terminus of the paragraph-2 underline relative to the following 。 is marginal at scan resolution; both transcribers and the question 62 quotation exclude the 。, which was adopted.
- 問題12: The 2-em dash was encoded as two U+2015 characters by both transcribers; the printed glyph is a single long dash, so a different but equally defensible encoding (e.g. U+2014 ×2) exists.
- 問題12: The booklet prints no source citation after the passage; this may be an omission in the retypeset source, but the transcription faithfully reflects the printed page.
- 問題13: The half-width space after circled digits ①-④ reflects a clearly visible printed gap, but it could be glyph side-bearing rather than a true typed space; rendering is unaffected either way.
- 問題13: The tab label マスダ買い取りサービス is printed in bold italic with an underline rule; italic and the rule cannot be represented in the allowed tag set, so only <b> is kept.
- 問題13: <th> vs <td> for the row-label cells is a semantic judgment based on the double rule setting off the label column; the labels themselves are printed in the same plain weight as body text.
Source-text quirks the OCR transcribed verbatim (printed typos/recall artifacts; the ones judged to be retype errors were corrected in the layer below) — 17 noted at OCR time:
- 問題1: Question 5 option 2 reads になって (likely the reading of 担って used as a distractor); verified against both image halves.
- 問題2: Question 13: a small stray hyphen-like mark is printed inside the blank, immediately before the closing parenthesis (looks like ( -)). Both transcribers and the adjudicator confirmed it in the double-resolution scan; it appears to be a retypesetting or scan artifact rather than intended content, so the blank was transcribed as the standard ( ) per convention.
- 問題3: Question 16: the underline covers 肝心な including な (consistent with the な-adjective options); verified at high zoom.
- 問題3: Question 17: the printed underline extends a full character cell under the sentence-final 。, so the period is included inside <u>…</u>; verified at high zoom (this typesetter otherwise ends underlines precisely, e.g. 14/16/18/19).
- 問題5: Question 31 option 3 is printed as 「買ったかにかかわらず」 (not 「買ったにもかかわらず」; confirmed independently by both transcribers and re-verified against the double-resolution image during adjudication); possibly a retypesetting artifact, but transcribed exactly as printed.
- 問題6: Question 36: the ★ occupies the first of the four slots as printed (印刷の ★ _ _ _), unlike the other questions where it is the third slot. Verified against the page image.
- 問題7: In the second paragraph, the printed text has a full-width space after 「言いたいのかね?」 before 「という声が聞こえる」; it is transcribed as printed.
- 問題7: The furigana over 雁 in 「月に雁」 is printed as かりがね (not がん or かり); verified against the double-resolution scan at 4x zoom.
- 問題8: Passage (2): the email header is printed as 宛て先/件 名/日 時 label:value lines between double horizontal rules; rendered as a <table> per convention, dropping the printed :separators (replaced by the cell boundary) and the full-width justification spaces inside the labels 件名 and 日時. The label is printed 宛て先 (with て), transcribed as printed. The double rule that ends the email after 担当:上田 映子 is likewise not representable and omitted. The lead-in sentence 「以下は、ある電気店から届いたメールである。」 is printed above the ruled block and included as the first <p>.
- 問題8: Passage (2): the time is printed 13:30 with a full-width colon; normalized to half-width 13:30 per convention, with digits normalized to half-width ASCII. Other printed full-width colons (e.g. 担当:) are kept as printed.
- 問題9: Passage (3) paragraph 4 prints a full-width space after both ? and ! (「君、広報が好きなの? じゃあ、やってみなさい! 自分で言うなら…」); transcribed as printed.
- 問題10: The passage prints 「シュバイッツア博士」 (small ッ followed by full-size ツ and full-size ア, no small ァ or long vowel mark) for Eugene Smith's Schweitzer photo; transcribed as printed even though シュバイツァー would be the usual spelling.
- 問題10: 「多くの人たちによさが認めてもらえる環境が整った」 is printed with が after よさ (one might expect よさを); verified at high resolution; possibly a source typo but transcribed as printed.
- 問題11: Passage A, first paragraph: printed as 「書こうとしている人やすでに」 with no comma after 人や — transcribed as printed.
- 問題11: Passage A, second body paragraph: 「…を信じる律儀さと同じで、たしかに真面目で素直ないい人…」 — the phrasing 人は…律儀さと同じで is slightly awkward (one might expect 信じる人と同じで) but it is transcribed exactly as printed; 素直ないい人 (two い) confirmed at high zoom during adjudication (transcriber B had read a single い).
- 問題12: No source citation line (e.g. (○○『…』による)) is printed after the passage; the 注 glosses follow the final paragraph directly. Unusual for an N1 問題12 passage and possibly an omission in this retypeset.
- 問題12: The passage uses both 面倒がないところに変化はない and 面倒のない人間関係 (が vs の) — transcribed as printed and verified against the page image.
Answer Confidence
65 high · 0 medium · 1 low · 0 unresolved (of 66).
Generated from answer-audit.json at build time. Questions listed below are below high confidence or have source disagreement.
Answer Sources
jlptzhen(key) — https://www.jlptzhen.com/n1%E7%9C%9F%E9%A2%98%E5%9C%A8%E7%BA%BF%E5%81%9A2023%E5%B9%B412%E6%9C%88%E6%97%A5%E6%9C%AC%E8%AF%AD%E8%83%BD%E5%8A%9B%E8%AF%95%E9%AA%8C/sohu (精英者教育 / 苏曼日语, teacher recall version)(key) — https://www.sohu.com/a/741118539_121124309luyenthitiengnhat.edu.vn(key) — https://www.luyenthitiengnhat.edu.vn/mod/page/view.php?id=3531&lang=jayuuki-bui-toi-yeu-ngoai-ngu(key) — N1_2023-12_answers.pdfjlpt-freq/extracted_text/N1/2023/2023.12_Answer N1 T12-2023.txt(key) — github.com/5Mcv4He/jlpt-freq extracted_text/N1/2023/2023.12_Answer N1 T12-2023.txtsolver-1(solver)solver-2(solver)manual-edits(manual)
Questions Needing Attention
| Question | Chosen | Confidence | Notes | Votes |
|---|---|---|---|---|
| q23 | 3 | high | jlptzhen=3, sohu=3, luyenthitiengnhat.edu.vn=2, yuuki-bui-toi-yeu-ngoai-ngu=3, jlpt-freq/text/N1/2023/2023.12_Answer N1 T12-2023=3, solver-1=3, solver-2=3, manual-edits=3 | |
| q39 | 3 | high | sohu=3, luyenthitiengnhat.edu.vn=3, yuuki-bui-toi-yeu-ngoai-ngu=3, jlpt-freq/text/N1/2023/2023.12_Answer N1 T12-2023=3, solver-1=3, solver-2=3, manual-edits=4 | |
| q63 | 4 | low | luyenthitiengnhat.edu.vn=4, yuuki-bui-toi-yeu-ngoai-ngu=3, jlpt-freq/text/N1/2023/2023.12_Answer N1 T12-2023=3, solver-1=4, solver-2=4 |