Job Search

Recruit Detail

Find out more about the work we do, the experience and skills we can bring to the table, and our terms and conditions.

Company Name

Stockmark Co., Ltd.

Job Type

1012 [Dev] VLM / MLLM Researcher

Work Detail

[Job Description] You will be responsible for research and development of VLM/Document AI targeting Japanese business and manufacturing documents within the Research Division's VLM unit. This includes not only technical development tasks such as model development, data synthesis, and evaluation, but also writing papers and submitting them to international conferences. ■Specific Job Responsibilities - Research and development of VLM/MLLM specialized for business documents and manufacturing - Model improvement related to document reading, table understanding, chart understanding, drawing analysis, chemical formula understanding, and layout understanding - Research and development related to VQA, CoT, Visual Grounding, Document Parsing, and Document Structuring - Construction of training data, evaluation data, benchmarks, annotation criteria, and quality filtering methods - Research on reliability evaluation of VLM output, hallucination suppression, evidence presentation, and validation methods - External communication such as writing papers, presenting at academic conferences, writing technical blogs, and publishing models and datasets Currently, we are conducting collaborative research with Professor Inui of Tohoku University and Researcher Ishigaki of AIST. *Scope of changes: All development tasks [Assigned Team] VLM Unit (Currently 4 members) [Organizational Structure] Research Division (8 members) - Knowledge Unit (4 members) - VLM Unit (4 members) [Main Programming Languages ​​Used] Python, SQL

Ideal Profile

Required Skills *All are required - Deep knowledge and work experience in image processing and natural language processing (especially practical experience in image processing and natural language processing using deep learning) - Proven track record of paper acceptance at peer-reviewed international conferences and in English-language academic journals - Graduate degree in science and engineering Preferred Skills 1. Research and practical experience in the following technical areas 1. Research or practical experience in VLM / MLLM / Document AI / OCR / Layout Analysis 2. Experience in using, evaluating, and improving related models such as Qwen-VL, InternVL, LLaVA, LayoutLM, etc. 3. Research and development experience in document reading, chart understanding, and drawing analysis 4. Experience in model improvement using SFT, Instruction Tuning, GRPO, DPO, RLHF, etc. 5. Domain knowledge in manufacturing, chemistry, machinery, electrical engineering, materials science, patents, and research and development 2. Business-level English proficiency Desired Candidate Profile - Inherently interested in technology and the ambition to become a leading expert in that field - I enjoy thinking about how to deliver to customers and grow the business while communicating with the business side. - I like following and verifying the latest theories by reading research papers. - Must be able to reside in Japan. - Must be able to communicate smoothly in Japanese, including technical matters related to development. Target: Business conversation level

Work Location

[Head Office Location] ■ LIFORK MINAMI AOYAMA S209, 1-12-3 Minami Aoyama, Minato-ku, Tokyo [Work Location] ■ Head office or home or any other location where remote work is possible (no restrictions on changes) *Full remote work is permitted

Phd. Stating Salary

■Expected Annual Salary: ¥7,500,000 - ¥15,000,000 *Monthly salary x 12 months ■Monthly Salary: ¥625,000 - ¥1,250,000 Basic Salary: ¥408,423 - ¥871,847 Life Planning Allowance: ¥55,000 Fixed Overtime Pay for 45 hours: ¥161,577 - ¥323,153 (Any excess will be paid separately) ■Salary Increase (Twice a year / May, November)

Similar Recruits

CyberAgent, Inc.

Job Type
[AI Lab] Research Internship (for doctoral students, theme selection type)

[Overview] AI Lab's research internship is exclusively for doctoral students. Working alongside active research scientists, interns will tackle various technological challenges at AI Lab, focusing on practical and advanced research themes utilizing AI technology. Research results can be submitted to and accepted at top international conferences in various academic fields; if accepted, all conference participation costs will be covered by the company. The timing, working hours, and format are negotiable, making it easier for those unable to commit fully for two months or those living far away, including overseas, to apply (for those living far away who choose to commute to the office, the company has a track record of covering accommodation (including communication costs)). [Examples of Research Themes] User understanding and robot behavior learning using real-world interaction data / Research on next-generation HCI media that extends real-world experiences Situational understanding and decision-making of autonomous mobile robots in store environments / Verbalization and analysis of real-world data / Multimodal platform models for understanding human behavior Research on graphic design production and editing Research on understanding the structure of manga and creative support technologies using LLM, VLM, and Agent Generation and understanding of internet advertisements using natural language processing / Research on output quality of large-scale language models Research on bandit algorithms for online advertising delivery Black box optimization <3D Vision & Graphics> 3D vision and computer graphics to support content creation Research on video production and editing Video Comprehension Support on AbemaTV [Achievements] Interns have had a total of 67 papers accepted (44 of which are from peer-reviewed international conferences and journals, including top conferences such as CVPR, ICML, IUI, ICCV, and NAACL).