Nishika Open-Sources J-MeetEval For Japanese Meeting Summaries
Nishika released J-MeetEval, a benchmark for instruction following in Japanese meeting summarization, as open source on GitHub and Hugging Face. The dataset holds 197 synthetic meeting transcripts, 52 instruction definitions, 408 total instructions and 27 instruction types across 16 categories, with inputs of 3,000 to 11,000 characters. Scoring is binary per instruction, judged by gpt-5-mini with a three-vote majority. Across seven models in the 4B to 9B class, scores on the Nejumi Leaderboard ran opposite to pass rates here (Spearman rho of minus 0.52).