AI Evaluation Scientist / English【JAPAN AI採用】

  • 正社員
  • 1000万
この求人の紹介
を受けたい

募集要項

仕事内容

About JAPAN AI

JAPAN AI, Inc. was established in April 2023 as a group company of Geniee, Inc. (TSE Growth Market) with the mission of dramatically expanding human potential through AI technology. We drive cutting-edge AI R&D both domestically and internationally.


Why We're Hiring

JAPAN AI is rapidly expanding its enterprise AI agent suite, including JAPAN AI AGENT / CHAT / SPEECH. As the core of our products shifts to LLMs and multi-agent systems, we are establishing a new specialized organization to scientifically evaluate the quality, safety, and reliability of AI outputs.


Mission

"Make AI Output Quality a ScienceーProve Agent Reliability through Research and Development of Evaluation Methods."

You will quantitatively evaluate and improve the output quality of LLMs and AI agents using methods from machine learning, statistics, and psychometrics. This position is not for "people who test"ーit is for "scientists who define and measure what makes a good AI."


Role&Expectations

As an AI Evaluation Scientist, you will lead the design, construction, and operation of the AI agent quality-evaluation infrastructure.


Research and develop evaluation metricsーscientifically define "what constitutes quality" through LLM-as-Judge calibration, reward modeling, and benchmark design

Design and build automated evaluation pipelinesーintegrate research outcomes into production CI/CD to deliver scalable quality gates

Red teaming and safety verificationーautomate adversarial testing and build policy compliance verification frameworks

Drive quality improvement through statistical experimental designーquantitatively verify the effectiveness of prompt strategies and model changes through A/B tests and significance testing

Feed evaluation signals back to research and development teamsーbuild a compound-interest loop for model improvement

Ensure the quality of products used in production by~200 companies through a "science of quality" approach


Why You'll Love This Role

Evaluation Science in practice : Practice "AI Evaluation Science"ーthe discipline that Apple, Anthropic, Scale AI, and others are investing inーwithin the context of Japanese enterprise AI. This is a globally rare position where evaluation methodology itself is the research subject.

A new application of ML/DS skills : Apply your machine learning and statistics expertise not to "building models" but to "evaluating models." Intellectual challenges span both research and implementationーreward modeling, LLM-as-Judge calibration theory, and benchmark design.

Quality determines product trust : In a production environment used by~200 companies, the evaluation infrastructure you build becomes the last line of defense for release quality. You will feel the direct business impact of quality assurance.

Greenfield position : Design and build the entirely new specialized domain of AI agent evaluation science from scratch. You will have significant autonomyーfrom evaluation metric R&D to production deployment of automated evaluation pipelines.

Frontline of AI safety : Engage in Responsible AI practices including automated red teaming, adversarial testing, and policy compliance verification. You will play a key role in scientifically guaranteeing safety in a world where AI agents autonomously execute business operations as "the brain of the enterprise."

Rapid-growth environment : In a startup that has grown to 200+people and 9 products in just 3 years, you will have significant autonomy in technical decision-making. You will work closely with Research Engineers and Agent Harness Engineers, influencing quality across the entire product suite.


Job Description

As an AI Evaluation Scientist, you will lead the design, construction, and operation of the AI agent Evaluation Infrastructure.


Evaluation Metric Research&Development

Research and implement LLM-as-Judge calibration methods (rubric design, bias detection, proper scoring rules)

Design, build, and validate evaluation benchmarks (construct validity, contamination detection)

Research the application of reward modeling / preference learning to evaluation

Select and design evaluation metrics (win rate, task success, factuality, harm detection)

Design, build, and maintain evaluation sets (synthetic data+real logs)

Automated Evaluation Pipeline Design&Development

Design and implement scalable automated evaluation pipelines

Integrate evaluation pipelines into CI/CD and build quality gates

Design agent evaluation harnesses (multi-turn, tool use, long-context support)

Ensure reproducibility and reliability of evaluation pipelines

Safety&Quality Verification

Research and implement automated red teaming (automated adversarial testing)

Build safety and policy compliance verification frameworks

Research and implement hallucination detection and calibration methods

Design and execute prompt / tool regression tests

Statistical Analysis&Experimental Design

Design and analyze statistical experiments (A/B tests, significance testing)

Visualize quality trends and automate regression detection

Create quality reports and improvement proposals

Feed evaluation signals back to research and development teams


Key Results (KR/Metrics)

Evaluation coverage rate (test case coverage)

Regression detection rate (pre-release quality degradation detection≧95%)

Evaluation pipeline execution time (completed within CI/CD)

LLM-as-Judge and human evaluation agreement rate

False positive / false negative rate

Safety incident rate (post-release)


Team Structure

Approximately 120 members are part of the development organization.

The AI Evaluation Scientist operates as a dedicated quality assurance function, collaborating closely with:


Agentic Product EngineerーAgent feature development

Research EngineerーResearch and development, model improvement

Agent Harness Engineer / Software Engineer (AI Platform)ーAI execution infrastructure development

Product ManagerーProduct design and quality requirements definition

経験・資格
※求人情報の応募要件全てに該当しなくても、企業様に対して内々に打診したり相談することが可能な場合もございます。一つでも当てはまる方は前向きにご検討下さい。

You May Be a Good Fit If You

Education&Experience

Master's degree or higher (or equivalent practical experience) in Computer Science, Machine Learning, Statistics, Mathematics, Physics, Psychometrics, or related fields

Practical experience as an ML Engineer, Data Scientist, Research Engineer, or in ML/AI evaluation-related roles

Technical Skills

Deep knowledge of LLM / generative AI evaluation methods (benchmark design, LLM-as-Judge, quantitative output quality measurement, hallucination detection, etc.)

Practical knowledge of statistics and experimental design (hypothesis testing, A/B testing, confidence intervals, effect sizes, etc.)

Experience building ML / evaluation pipelines in Python

Practical experience with machine learning frameworks (PyTorch, JAX, TensorFlow, etc.)

Experience designing and implementing evaluation metrics (task-specific metric design beyond precision/recall)

Language requirement (at least one of the following):

Japanese: Fluentーable to discuss product development without friction

English: Business level


This position is a research and development role responsible for AI output Evaluation Science. Research or implementation experience in ML model evaluation / LLM evaluation is required.


Strong Candidates May Also Have

Publication experience at top ML/NLP conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, etc.)

Research or implementation experience with reward modeling / preference learning (RLHF, DPO, etc.)

Experience with LLM-as-Judge calibration and rubric design

Knowledge or experience in AI safety, Responsible AI, and red teaming

Experience with benchmark design and validity verification (IRT, construct validity)

Experience evaluating multi-agent workflows, tool use, and long-context scenarios

Large-scale data processing experience (Spark / BigQuery, etc.)

Experience integrating ML / evaluation pipelines into CI/CD

Ability to read, comprehend, and reproduce research papers

Technical communication ability in English


Tech Stack

Languages : Python (evaluation pipelines&analysis) , TypeScript / React / Next.js (frontend) / NX

Evaluation/QA : pytest, LangSmith, Weights&Biases, custom eval frameworks

Data : BigQuery, Spark, Pandas

Infrastructure : GCP (containers / K8s) , Docker, Terraform

CI/CD : GitHub Actions

Tools : Slack, Confluence, Linear, Google Workspace, GitHub, Notion

AI Dev Support: Claude Code MAX Plan, Cursor, ChatGPT, Devin

Work environment : Mac (Apple Silicon) , dual monitors available

※更なる詳細事項はカウンセリング(面談)時にお伝えします。
勤務地
東京都新宿区西新宿6-8-1 住友不動産新宿オークタワー5/6階
【受動喫煙対策の有無:有】
敷地内禁煙(屋外に喫煙場所設置)
想定年収
800 万円 ~ 1600 万円
福利厚生・待遇
【待遇・福利厚生】
<正社員>
・書籍購入補助(半期 30,000円まで)
・リフレッシュ手当(毎月 5,000円まで)
・部活動手当(毎月5,000円まで)
・家賃手当(当社指定の駅を対象とし毎月30,000円まで)
・シャッフルランチ/ディナー(四半期に一度ランチ1,000円まで、ディナー5,000円まで)
・資格取得支援制度、英語学習支援制度(業務に必要な場合のみ)
・リフレッシュ休暇制度(3年間継続勤務した社員へ毎年付与される特別休暇 2日)
・定期健康診断(年1回)
・従業員持株会

<契約社員>
・書籍購入補助(半期 30,000円まで)
・リフレッシュ手当(毎月 5,000円まで)
・部活動手当(毎月5,000円まで)
・シャッフルランチ/ディナー(四半期に一度ランチ1,000円まで、ディナー5,000円まで)
・リフレッシュ休暇制度(3年間継続勤務した社員へ毎年付与される特別休暇 2日)
・定期健康診断(年1回)

【諸手当】
・交通費全額支給
勤務条件

勤務時間

10:00~19:00
※土日祝は休業日となります
※出向の場合は、出向先の規程に準じます
Work Style
Hybrid work : 3 days in office, 2 days remote
Flexible working hours : Core time is negotiable
Flexibility : Future consideration for more flexible work styles is possible

休日・休暇

完全週休二日制
所定休日:土・日・祝日
休暇:年次有給休暇、夏季休暇(3日)、年末年始休暇(12月31日~1月3日)、慶弔休暇

試用期間

1か月

加入保険

社会保険完備(健康保険:関東ITソフトウェア健康保険組合)

Recruiting No.
01008655000650

企業情報

社名
株式会社ジーニー
事業内容
広告プラットフォーム事業/マーケティングSaaS事業/海外事業/デジタルPR事業

エリートネットワーク取材班による独自解説

広告プラットフォーム事業を中心に、企業のデジタルマーケティングを支援するSaaS事業を展開。テクノロジー企業を標榜し、生成AIを使ったサービスを手掛けるJAPAN AI株式会社を2023年に立ち上げたほか、北米の大手広告テクノロジー企業Zeltoを子会社化するなど事業拡大を図っている。
創業6年で国内トップクラス規模に拡大したアドプラットフォームを有し、DSPやDMP、マーケティングオートメーション領域についても、順調にシェアを伸ばしている。DSPは広告500社、SSPはメディア20000社ほどあり、業界No.1の地位を固くしている。
Web広告などで培ったアドテクノロジーのノウハウを活かし、DOOH(Digital Out of Home)という“屋外広告 × デジタル × データ活用”の世界に参入。これにより、ただの看板売りではなく、テック × データ × 広告のクロス領域での強みを持っている。

蓄積してきたデータを活かしたマーケティングSaaS事業も好調で、CRMの領域でシェアを伸ばしてきている。今後は海外展開を含め、さらに伸ばしていく方針。
エンジニアを内製化しているため、技術力の高さが売り。
続きを読む

転職支援サービスの流れ

  • STEP1
    転職支援サービスへの
    お申し込み
    登録フォームに必要事項を入力のうえ、お申し込みください。

    既に応募書類を作成済の方はメールにてこちらから登録いただけます。
  • STEP2
    サービス利用開始の
    ご連絡
    お申し込みから3~4営業日以内に、電話かメールにて連絡をいたします。
    ※求人状況によっては、転職カウンセラーとの面談・相談サービスの提供が難しい場合もございます。
  • STEP3
    転職カウンセリング
    対面(オンライン可)または電話にて、あなたのご要望やご経験、今回の転職で実現したい事柄の優先順位などをお伺いします。
  • STEP4
    求人紹介・応募書類
    チェック
    様々な角度から具体的な企業や求人のご提案をいたします。各社の社風や雰囲気についてもご案内します。
    応募の意向をお知らせください。応募書類一式の添削も承ります。
  • STEP5
    企業への推薦・書類選考
    応募先が決まりましたら、転職カウンセリングにてお伺いした内容をもとに、書類選考を通過すべく作成した魅力的な「推薦状」を添えて、各企業様に打診いたします。
  • STEP6
    面接
    面接の日程調整は転職カウンセラーにお任せください。
    ご希望により、面接に臨むにあたっての事前準備や、各企業毎の面接対策も承っております。
  • STEP7
    内定・退職フォロー
    最終選考に合格した後も、給与や待遇についての交渉は遠慮なく転職カウンセラーにご相談ください。入社企業が決まりましたら現職の円満退社に向けたフォローも承ります。
  • STEP8
    入社
    入社時期のすり合わせ、諸手続き等、入社に至るまで転職カウンセラーがサポートいたします。
    入社後は新天地にて思う存分ご活躍ください。
    ※エリートネットワークを通じて転職に成功された方々の『転職体験記』はこちらからご覧ください。

この求人を見た方に
おすすめの求人

    条件を変えて検索する

    希望職種
    希望業界
    希望勤務地
    フリーワード
    希望年収
    推奨年齢
    こだわり条件

    ご希望の職種を選択してください

    ご希望の職種を選択してください

    ご希望の業界を選択してください

    ご希望の勤務地を選択してください

    ご希望条件を入力ください

    希望年収
    推奨年齢
    フリーワード

    こだわり条件を入力ください

    転職体験記

    • TOEIC960点の25歳。コンサル会社での金融システム開発を活かし、より社会貢献実感を持てるプライム上場 日本最大級の発電会社へ転職
      25歳/都立高校 卒 一橋大学 経済学部 卒 TOEIC 960点 日商簿記検定2級 ...

      【東証プライム上場 日本最大級の発電会社】 電力需給統括部 電力需給システム開発プロジェクトリード担当

      転職体験記を読む
    • 大学院でのニューラルネットワークの学びが活かせない通信会社の理系修士28歳。業務改善×データ活用の知識で、老舗日系コンサル会社へ
      28歳/都立高校 卒 私立大学 理工学部 電気電子工学科 卒 私立大学大学院 理工...

      【総合重電機メーカー直系 老舗コンサルティング会社】 先端テクノロジー・データサイエンス分野のコンサルタント

      転職体験記を読む
    • 「1級陸上無線技術士」を持つ放送エンジニア。34歳で抱いた危機感を機に、特殊な業界経験の言語化に苦戦するも、自分一人では届かないと諦めていた世界的な大手電機メーカーグループの技術系総合職へ。
      34歳/地方県立高校 卒 都内有名私立理工系大学 工学部 電子工学科 卒 地方国立...

      【東証プライム上場 有名電機メーカーグループ】 技術系総合職

      転職体験記を読む

    注目企業インタビュー

    • 全員参加型のビジネス変革が成果を生み出し、キャリア人材の成長機会が増え続けています。

      富士通株式会社
      CHRO室 シニアディレクター 黒川 和真氏
    • 人々の生活や命を支えるため、「食料・水・環境」分野で地域に根ざした事業にチャレンジする

      株式会社クボタ
      人事部 採用室長 猪野 陽一氏
    • 高度な専門性を持ち、お客様の業務に精通したSEと営業が一丸となり、 お客様のビジネスの成長を “攻めと守り”のITで支援。

      新日鉄住金ソリューションズ株式会社
      エグゼクティブプロフェッショナル 人事本部 キャリア採用センター所長 岡田 康裕氏
    • 世界に向かうデジタルビジネスのパートナーとして、売上拡大とコスト最適化を支援しています。

      トランスコスモス株式会社
      執行役員 人事本部担当 兼 サービス推進本部人材マネジメント統括部担当 名倉 英紀氏
    • エネルギー、インフラ、ストレージ。3つの注力事業において、新しい人材が 「新生東芝」 を動かし始めています。

      株式会社東芝
      人事・総務部 人材採用センター センター長 三橋 一仁氏
    • グローバル展開する企業のプライムパートナーとして、経営から製造現場まで、多様な課題の解決をITで支援。

      東洋ビジネスエンジニアリング株式会社
      ソリューション事業本部 事業本部長 取締役 別納 成明 氏
      企業インタビュー一覧はこちら
    • 富士通株式会社

      全員参加型のビジネス変革が成果を生み出し、キャリア人材の成長機会が増え続けています。

    • 株式会社クボタ

      人々の生活や命を支えるため、「食料・水・環境」分野で地域に根ざした事業にチャレンジする

    • 新日鉄住金ソリューションズ株式会社

      高度な専門性を持ち、お客様の業務に精通したSEと営業が一丸となり、 お客様のビジネスの成長を “攻めと守り”のITで支援。

    • トランスコスモス株式会社

      世界に向かうデジタルビジネスのパートナーとして、売上拡大とコスト最適化を支援しています。

    • 株式会社東芝

      エネルギー、インフラ、ストレージ。3つの注力事業において、新しい人材が 「新生東芝」 を動かし始めています。

    • 東洋ビジネスエンジニアリング株式会社

      グローバル展開する企業のプライムパートナーとして、経営から製造現場まで、多様な課題の解決をITで支援。

       企業インタビュー一覧はこちら
    この求人の紹介を受けたい お気に入り
    この求人の紹介を受けたい お気に入り