Nishika Open-Sources J-MeetEval For Japanese Meeting Summaries
General-purpose leaderboard rank does not predict whether a Japanese model will follow formatting and scope instructions on long meeting transcripts, which is the gap Nishika built a task-specific benchmark to measure.
Reporting from 1 source: ASCII.jp.
Nishika released J-MeetEval, a benchmark for instruction following in Japanese meeting summarization, as open source on GitHub and Hugging Face. The dataset holds 197 synthetic meeting transcripts, 52 instruction definitions, 408 total instructions and 27 instruction types across 16 categories, with inputs of 3,000 to 11,000 characters. Scoring is binary per instruction, judged by gpt-5-mini with a three-vote majority. Across seven models in the 4B to 9B class, scores on the Nejumi Leaderboard ran opposite to pass rates here (Spearman rho of minus 0.52).
J-MeetEval sets several instructions at once on a single transcript, such as unifying numerals to full-width characters, sorting by speaker order, and returning a fixed JSON shape, then scores each one pass or fail. Formatting results split by direction: rounding numerals to half-width passed 98 percent of the time, while the opposite demand for full-width managed 27 percent. Instructions that narrow scope were harder still, with extracting only decisions and their supporting remarks at 25 percent. Combining instructions also hurt, with chronological ordering dropping sharply when paired with anonymization, job titles or JSON output.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.