Hi there Welcome to my Homepage!

I am a master’s student in Software Engineering at Jilin University and a research assistant at the Trust & Application AI Lab (TAI Lab), HKUST(GZ). My research focuses on trustworthy large language model systems, with recent work on LLM router security, black-box model fingerprinting, post-training data extraction from diffusion language models, and realistic benchmarks for protein modeling.

Feel free to reach out if you are interested in collaboration or potential research opportunities.

News

  • 2026.04 R2A was accepted to ACL 2026 Main Conference and released on arXiv.
  • 2026.03 Protap, a benchmark for realistic protein modeling applications, was updated on arXiv.
  • 2025.05 DuFFin was released on arXiv.
  • 2024.09 I joined TAI Lab at HKUST(GZ) as a research assistant.

Experience

Trust & Application AI Lab, HKUST(GZ)
2024.09 - Present
Research Assistant at TAI Lab, led by Prof. Enyan Dai.
Research on emerging threats in LLM services, including model IP protection and LLM router security.
Jilin University
2024.09 - Present (Expected 2027.09)
M.S. student in Software Engineering, recommended admission.
Ongoing research on trustworthy AI systems and LLM security.
Jilin University
2020.09 - 2024.09
B.E. in Software Engineering.
Undergraduate training in software engineering and AI systems.

Publications

(* equal contribution · † corresponding author · ‡ project leader)

DuFFin paper thumbnail
DuFFin: A Dual-Level Fingerprinting Framework for LLMs IP Protection
Yuliang Yan, Haochun Tang, Shuo Yan, Enyan Dai
DuFFin verifies LLM ownership in black-box settings through trigger-response and knowledge-level fingerprints, remaining effective under fine-tuning, quantization, and safety alignment.
EACL 2026   [arXiv] [PDF] [code]
Protap paper thumbnail
General Protein Pretraining or Domain-Specific Designs? Benchmarking Protein Modeling on Realistic Applications
Shuo Yan, Yuliang Yan, Bin Ma, Chenao Li, Haochun Tang, Jiahua Lu, Minhua Lin, Yuyuan Feng, Enyan Dai
Protap systematically compares general pretraining and domain-specific protein modeling designs on realistic downstream applications.
KDD 2026   [arXiv] [PDF] [code]
D-Miner: Extracting Post-Training Data from Diffusion Language Models
Haochun Tang
D-Miner starts from pure mask inputs, identifies member-like candidates through reconstruction signals, converts the signal into hidden-state guidance, and steers denoising toward post-training data regions.
Under Review  

Projects

StudyClawHub project thumbnail
StudyClawHub
I participate in the development and maintenance of StudyClawHub, an AI Agent and Skill ecosystem for students. The platform supports reusable agent browsing, task skill installation, custom skill submission, and GitHub-driven community publishing.
Platform Development   [project]