Entry 262
The tree cleaned of the games' dialogue
Filesplans/release.mdtools/test_stream_rs.pycrates/piney-stream/tests/subtitle_fixture.txt
The repository is to be public, under MIT OR Apache-2.0. It must not carry the games' writing. The policy, in plans/release.md:
- Never: dialogue and writing in any form, whether plain, hex, in fixtures or as test input.
- Allowed where the work needs them: names and labels. That means area words, item and character names, menu labels, the menus' short system messages, and asset and symbol names.
How it was found
The corpus is every printable run of 16 characters or more with two spaces
or more, from every file of the four discs. DATA.BIN and STREAM/*.BIN
members were inflated, and movies and raw audio skipped: 27,159 runs.
Every 24-character window at a word start that reads like a sentence is
hashed. Each tracked file's own windows are looked up in that set, and so
are the windows of its hex runs decoded.
The first pass found 207 matches in 69 files. A second pass at 16-character windows found the short quotes the first missed. A sweep for Japanese characters covered what an ASCII corpus cannot see.
What was changed
- The subtitle fixture (
crates/piney-stream/tests/subtitle_fixture.txt) held about 100 cut-scene lines, hex encoded in itschangerows. The Rust test never read those rows, since it compares only the draws' hashes. The rows now carry each string's FNV-1a 64 (401 strings), and the generator writes them that way. - Quotes in prose. Dialogue quoted in docs, code comments and worklogs
became a description of what the line is. Worklog entries 18, 70, 87, 92,
123, 134, 143 and 228 were redacted this way. It is the only edit made to
past entries, and this entry records it. The fountain's conversation
(
menus/fountain.rs,docs/engine/field-ui.md) now goes by message numbers. - Sample text in the harnesses. The harnesses that feed text to the
game's code and to the port alike (
test_desktop_rs.py,test_chat_msg_rs.py,test_fieldui_rs.py) now use neutral lines of the same lengths.test_evscript.pychecks that a line is inlined with its voice without spelling the line out.
A finding on the way
The chat-balloon harness failed on a 74-byte sample line. The game's
ccChatMsg::OpenChat (INF main 0x001a67c0) copies the text with a plain
strcpy into data[i], which is 71 bytes. It then measures the source
with ccKanjiStrWidth, not the copy. A longer line overruns into the next
balloon's buffer. The port cuts the text at 70 bytes instead, so the two
differ only past the game's own buffer. The game's lines never get there:
the longest sample of the old set was 68 bytes. The new samples keep to
70.
Checks
- Rescan at 24-character windows: 82 matches left. Each is a name, a
system message, a debug string in an IOP module (
playEnv running status error), or ordinary English that happens to match ("in the fields and dungeons"). - Rescan at 16 characters: the same kinds, and no dialogue.
- The harnesses pass:
test_desktop_rs(message window),test_chat_msg_rs,test_fieldui_rs(field message) andtest_evscript. So do the Rust tests over the fixture:the_subtitles_are_the_games, the desktop'sevent_message_freezes_the_desktopand piney-audio's voice tests.
- Short dialogue under four words is not caught by the scan. - Text in pictures (the games' textures) is not text to it. - The history's blobs have not been scanned. The history holds whole data files, so it will not be published as it is (plans/release.md).