About JAPAN AI
JAPAN AI, Inc. was established in April 2023 as a group company of Geniee, Inc. (TSE Growth Market) with the mission of dramatically expanding human potential through AI technology. We drive cutting-edge AI R&D both domestically and internationally.
Why We're Hiring
JAPAN AI is rapidly expanding its enterprise AI agent suite, including JAPAN AI AGENT / CHAT / SPEECH. As the core of our products shifts to LLMs and multi-agent systems, we are establishing a new specialized organization to scientifically evaluate the quality, safety, and reliability of AI outputs.
Mission
"Make AI Output Quality a ScienceーProve Agent Reliability through Research and Development of Evaluation Methods."
You will quantitatively evaluate and improve the output quality of LLMs and AI agents using methods from machine learning, statistics, and psychometrics. This position is not for "people who test"ーit is for "scientists who define and measure what makes a good AI."
Role&Expectations
As an AI Evaluation Scientist, you will lead the design, construction, and operation of the AI agent quality-evaluation infrastructure.
Research and develop evaluation metricsーscientifically define "what constitutes quality" through LLM-as-Judge calibration, reward modeling, and benchmark design
Design and build automated evaluation pipelinesーintegrate research outcomes into production CI/CD to deliver scalable quality gates
Red teaming and safety verificationーautomate adversarial testing and build policy compliance verification frameworks
Drive quality improvement through statistical experimental designーquantitatively verify the effectiveness of prompt strategies and model changes through A/B tests and significance testing
Feed evaluation signals back to research and development teamsーbuild a compound-interest loop for model improvement
Ensure the quality of products used in production by~200 companies through a "science of quality" approach
Why You'll Love This Role
Evaluation Science in practice : Practice "AI Evaluation Science"ーthe discipline that Apple, Anthropic, Scale AI, and others are investing inーwithin the context of Japanese enterprise AI. This is a globally rare position where evaluation methodology itself is the research subject.
A new application of ML/DS skills : Apply your machine learning and statistics expertise not to "building models" but to "evaluating models." Intellectual challenges span both research and implementationーreward modeling, LLM-as-Judge calibration theory, and benchmark design.
Quality determines product trust : In a production environment used by~200 companies, the evaluation infrastructure you build becomes the last line of defense for release quality. You will feel the direct business impact of quality assurance.
Greenfield position : Design and build the entirely new specialized domain of AI agent evaluation science from scratch. You will have significant autonomyーfrom evaluation metric R&D to production deployment of automated evaluation pipelines.
Frontline of AI safety : Engage in Responsible AI practices including automated red teaming, adversarial testing, and policy compliance verification. You will play a key role in scientifically guaranteeing safety in a world where AI agents autonomously execute business operations as "the brain of the enterprise."
Rapid-growth environment : In a startup that has grown to 200+people and 9 products in just 3 years, you will have significant autonomy in technical decision-making. You will work closely with Research Engineers and Agent Harness Engineers, influencing quality across the entire product suite.
Job Description
As an AI Evaluation Scientist, you will lead the design, construction, and operation of the AI agent Evaluation Infrastructure.
Evaluation Metric Research&Development
Research and implement LLM-as-Judge calibration methods (rubric design, bias detection, proper scoring rules)
Design, build, and validate evaluation benchmarks (construct validity, contamination detection)
Research the application of reward modeling / preference learning to evaluation
Select and design evaluation metrics (win rate, task success, factuality, harm detection)
Design, build, and maintain evaluation sets (synthetic data+real logs)
Automated Evaluation Pipeline Design&Development
Design and implement scalable automated evaluation pipelines
Integrate evaluation pipelines into CI/CD and build quality gates
Design agent evaluation harnesses (multi-turn, tool use, long-context support)
Ensure reproducibility and reliability of evaluation pipelines
Safety&Quality Verification
Research and implement automated red teaming (automated adversarial testing)
Build safety and policy compliance verification frameworks
Research and implement hallucination detection and calibration methods
Design and execute prompt / tool regression tests
Statistical Analysis&Experimental Design
Design and analyze statistical experiments (A/B tests, significance testing)
Visualize quality trends and automate regression detection
Create quality reports and improvement proposals
Feed evaluation signals back to research and development teams
Key Results (KR/Metrics)
Evaluation coverage rate (test case coverage)
Regression detection rate (pre-release quality degradation detection≧95%)
Evaluation pipeline execution time (completed within CI/CD)
LLM-as-Judge and human evaluation agreement rate
False positive / false negative rate
Safety incident rate (post-release)
Team Structure
Approximately 120 members are part of the development organization.
The AI Evaluation Scientist operates as a dedicated quality assurance function, collaborating closely with:
Agentic Product EngineerーAgent feature development
Research EngineerーResearch and development, model improvement
Agent Harness Engineer / Software Engineer (AI Platform)ーAI execution infrastructure development
Product ManagerーProduct design and quality requirements definition
You May Be a Good Fit If You
Education&Experience
Master's degree or higher (or equivalent practical experience) in Computer Science, Machine Learning, Statistics, Mathematics, Physics, Psychometrics, or related fields
Practical experience as an ML Engineer, Data Scientist, Research Engineer, or in ML/AI evaluation-related roles
Technical Skills
Deep knowledge of LLM / generative AI evaluation methods (benchmark design, LLM-as-Judge, quantitative output quality measurement, hallucination detection, etc.)
Practical knowledge of statistics and experimental design (hypothesis testing, A/B testing, confidence intervals, effect sizes, etc.)
Experience building ML / evaluation pipelines in Python
Practical experience with machine learning frameworks (PyTorch, JAX, TensorFlow, etc.)
Experience designing and implementing evaluation metrics (task-specific metric design beyond precision/recall)
Language requirement (at least one of the following):
Japanese: Fluentーable to discuss product development without friction
English: Business level
This position is a research and development role responsible for AI output Evaluation Science. Research or implementation experience in ML model evaluation / LLM evaluation is required.
Strong Candidates May Also Have
Publication experience at top ML/NLP conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, etc.)
Research or implementation experience with reward modeling / preference learning (RLHF, DPO, etc.)
Experience with LLM-as-Judge calibration and rubric design
Knowledge or experience in AI safety, Responsible AI, and red teaming
Experience with benchmark design and validity verification (IRT, construct validity)
Experience evaluating multi-agent workflows, tool use, and long-context scenarios
Large-scale data processing experience (Spark / BigQuery, etc.)
Experience integrating ML / evaluation pipelines into CI/CD
Ability to read, comprehend, and reproduce research papers
Technical communication ability in English
Tech Stack
Languages : Python (evaluation pipelines&analysis) , TypeScript / React / Next.js (frontend) / NX
Evaluation/QA : pytest, LangSmith, Weights&Biases, custom eval frameworks
Data : BigQuery, Spark, Pandas
Infrastructure : GCP (containers / K8s) , Docker, Terraform
CI/CD : GitHub Actions
Tools : Slack, Confluence, Linear, Google Workspace, GitHub, Notion
AI Dev Support: Claude Code MAX Plan, Cursor, ChatGPT, Devin
Work environment : Mac (Apple Silicon) , dual monitors available
10:00~19:00
※土日祝は休業日となります
※出向の場合は、出向先の規程に準じます
Work Style
Hybrid work : 3 days in office, 2 days remote
Flexible working hours : Core time is negotiable
Flexibility : Future consideration for more flexible work styles is possible
完全週休二日制
所定休日:土・日・祝日
休暇:年次有給休暇、夏季休暇(3日)、年末年始休暇(12月31日~1月3日)、慶弔休暇
1か月
社会保険完備(健康保険:関東ITソフトウェア健康保険組合)
【東証プライム上場 日本最大級の発電会社】 電力需給統括部 電力需給システム開発プロジェクトリード担当
【総合重電機メーカー直系 老舗コンサルティング会社】 先端テクノロジー・データサイエンス分野のコンサルタント
【東証プライム上場 有名電機メーカーグループ】 技術系総合職
全員参加型のビジネス変革が成果を生み出し、キャリア人材の成長機会が増え続けています。
人々の生活や命を支えるため、「食料・水・環境」分野で地域に根ざした事業にチャレンジする
高度な専門性を持ち、お客様の業務に精通したSEと営業が一丸となり、 お客様のビジネスの成長を “攻めと守り”のITで支援。
世界に向かうデジタルビジネスのパートナーとして、売上拡大とコスト最適化を支援しています。
エネルギー、インフラ、ストレージ。3つの注力事業において、新しい人材が 「新生東芝」 を動かし始めています。
グローバル展開する企業のプライムパートナーとして、経営から製造現場まで、多様な課題の解決をITで支援。