Skip to content
View xiaofengShi's full-sized avatar
🤔
Focusing
🤔
Focusing

Block or report xiaofengShi

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
xiaofengShi/README.md

Xiaofeng Shi

AI Researcher & Engineer at BAAI · Previously ByteDance and Meituan

I build foundation-model systems for specialized domains, from open data and post-training to retrieval agents and multimodal reasoning. My work connects research with reusable code, datasets, models, and evaluation tools.

Website · Google Scholar · Hugging Face · Email

Selected research

  • MechVQA / MechVL ICML 2026 — A benchmark and domain-specialized models for understanding mechanical drawings. Code · Data & models
  • IAR 2026 preprint — Staged post-training to internalize document knowledge, align question answering, and recover general capabilities without retrieval.
  • SPAR / SPARBench 2025 preprint — Multi-agent scholarly retrieval with an evaluation dataset. Code · Data
  • SciSage / SurveyScope 2025 preprint — Multi-agent scientific survey generation and evaluation. Code · Data

More work on post-training: Wnuan · RAFT · SFTKey · MoSLD (COLING 2025).

Multimodal retrieval and domain models: ChartWalker · CareBot (AAAI 2025) · Aquila-Med.

Full research index · Published patent applications

Open data at BAAI

Earlier open source

  • CHINESE-OCR CHINESE-OCR stars — Chinese scene-text detection and recognition; legacy project.
  • Image2Katex — Image-to-LaTeX recognition for printed and handwritten formulas.
  • DKT-TensorFlow — Deep knowledge tracing in TensorFlow.

Pinned Loading

  1. MechVQA MechVQA Public

    ICML 2026 benchmark and MechVL models for multimodal understanding of mechanical engineering drawings.

    Python 32 2

  2. SPAR SPAR Public

    SPAR: multi-agent scholarly retrieval with query decomposition, query evolution, and citation-aware exploration.

    Python 28 8

  3. FlagAI-Open/FlagAI FlagAI-Open/FlagAI Public

    FlagAI (Fast LArge-scale General AI models) is a fast, easy-to-use and extensible toolkit for large-scale model.

    Python 3.9k 417

  4. CHINESE-OCR CHINESE-OCR Public

    End-to-end Chinese scene-text detection and recognition with CTPN, CRNN, and CTC (legacy project).

    Python 3k 941

  5. handsomestWei/patent-disclosure-skill handsomestWei/patent-disclosure-skill Public

    中国专利.skill:专利点挖掘与交底书(发明/实用/外观)编写,通俗解读专利,嗅探政策动向,辅助审查答复。

    Python 9.3k 961

  6. Image2Katex Image2Katex Public

    Image-to-LaTeX recognition for printed and handwritten mathematical formulas.

    HTML 293 73