# 嘉永七年のアイヌ語について — research data

成田修一（2010）「嘉永七年のアイヌ語について」『二松学舎大学東アジア学術総合研究所集刊』40、25–69頁。
[Repository record](https://nishogakusha.repo.nii.ac.jp/records/2656) · [Sources bibliography](https://db.aynu.org/sources/2010-narita-kaei-nana-nen-ainu-go)

Version 1.1.0. The tables were transcribed by direct visual reading of the page images. No OCR was used. Full independent character proofreading remains pending. Each vocabulary row carries its own review status; `checked` describes the paper transcription, not the historical form’s linguistic interpretation.

## Files and reproduction

`transcription/38.psv` through `68.psv` are the editable appendix rows. Their five fields are Ainu kana, Japanese heading, source label, paper notes, and the 藻汐草 comparison. T=話通言, K=蝦夷紀行, S=唐太島記行, R=蝦夷地旅行記録, D=唐太日記. `@S` expands to 上欄外に「案ニ以下奥蝦夷詞歟」とある。

`comparisons.psv` contains printed page, section, sound feature, source abbreviation, Japanese heading, historical form, dictionary heading, and dictionary form. It records each occupied source cell and omits empty cells. It does not encode horizontal alignment between examples across sources. Dictionary romanizations reproduce the paper’s comparisons with 『アイヌ語方言辞典』.

`study.json` records bibliography, source identity and PDF checksum. `reviews.json` records visual checks and uncertainties, with SHA-256 digests of the five expanded data fields. A checked row or alignment whose content changes fails validation until reviewed again. `alignments.json` links checked rows to exact manuscript entry IDs and records spelling or heading differences. Run `bun run data:build` in the Records repository to produce `static/export/research/narita-2010/` and the site data. The CSV and JSON exports derive from the same inputs.

Downloads include the complete vocabulary, the 話通言 subset, comparison tables, the author’s reported counts, a count audit, unresolved readings, alignments, and BibTeX/CSL citations. `data.json` contains all tables and metadata. CSV files are UTF-8 with a BOM; review objects use JSON within a quoted CSV cell.

## Reading conventions

Historical spellings, duplicates and differences between columns are retained. Parenthetical ママ, question marks and □ generally reproduce the paper. `［判読保留］` marks an unresolved export reading; `［図形］` represents a small diagram. Ruby is flattened where transcribed, with incomplete coverage. Unicode 〱, ヽ, 々 and 〃 represent repetition marks; nonstandard handakuten uses combining U+309A.

The appendix has 1,213 transcribed rows, including four without an Ainu form. Its 1,209 rows with a form exceed the author’s p.27 total of 1,197. Five repeated form–gloss pairs are retained: the paper explicitly permits repeated occurrences (p.37), and the duplicate rows were checked in the images. Removing them would still leave 1,204. The two literal 蝦夷唐太日記 labels are unassigned. These measures remain unreconciled; no rows were removed or relabelled to force agreement.

The paper’s 蝦夷紀行 is attributed to 忠蔵. It is not linked by title to the Records work attributed to 間宮林蔵. The initial six entry links compare the later 話通言 evidence with 藻汐草. Textual dependence has not been excluded. The paper’s 未 heading for ビカタコバツ does not transfer to the 藻汐草 entry, where the small 未 accompanies ラレブニ.

## Citation and rights

Cite the paper with printed pages, for example **成田 2010: 58**. Identify the derivative transcription separately as *ERDAL, Narita 2010 research dataset, version 1.1.0*, with a row ID such as `narita2010-p58-r34`. PDF page 34 is navigation metadata for printed p.58. BibTeX and CSL files describe the paper; they do not attribute the derivative dataset to Narita.

The paper’s repository record does not identify an open reuse licence. Copyright in the paper remains with its rights holders. These exports contain factual table transcriptions with attribution; they do not include the PDF or manuscript images. See `study.json` for rights metadata.
