Đinh Thiên Ân
Software Engineer at VPBank · AI Engineer — LLMs, RAG & Agents, Vision-Language Models
Software Engineer at VPBank in Ho Chi Minh City and Computer Science graduate (B.Sc., GPA 3.5/4.0) of the University of Information Technology (UIT – VNU-HCM). Previously built an LLM + MCP test-automation agent as an AI Engineer Intern at TMA Solutions. Co-author of ZeroSemble (XLLM Workshop @ ACL 2025; 2nd place, DocIE shared task) and of two SemEval-2026 system papers (6th place in SemEval-2026 Task 4, Track B). Builds RAG, multi-agent, video-retrieval and vision-language systems; his team reached the final of the AI Challenge HCMC 2026 among 700+ teams.
Featured Projects
Research projects, competition entries, and production systems spanning NLP, computer vision, multi-agent systems, and full-stack development.
Zero-shot, two-stage pipeline for document-level entity and relation extraction. First, entities are extracted by three heterogeneous LLMs (DeepSeek-R1 distill, Llama-3.3-70B, Qwen-2.5-32B) and consolidated through deduplication and majority-vote typing. Then relations are extracted with prompts that constrain head and tail to the consolidated entity set. No domain-specific training is required.
Role:
Co-author (2nd author); repository owner
Highlights:
- 2nd place, DocIE shared task @ XLLM Workshop, ACL 2025 (team UIT-SHAMROCK)
- Published in ACL Anthology (2025.xllm-1.25)
- F1 scores: EI 55.65, EC 26.11, REG 4.19, RES 4.01, Overall 22.49
Interactive search system for finding events in a large video archive. The index covers 1,487 videos in 125 groups and 524,891 keyframes (news, CCTV traffic cameras, cycling). Features: Vietnamese text-to-image visual search on SigLIP2 with LLM query translation, OCR text search, ASR transcript search, object-detection filtering, temporal multi-event search, voice queries, and a built-in DRES submission bar.
Role:
Built the entire system
Highlights:
- Finalist, AIC 2026 (AI Challenge HCMC), 700+ teams
- Index: 1,487 videos · 125 groups · 524,891 keyframes
- Latency: OCR ~0.17s, ASR ~0.11s, object detection <1s, temporal search ~4s
Searches, gathers and analyzes online information (news, social media, public data) about customers from their personal or business identification details. It builds a risk blacklist from crawled news with LLM-based extraction and risk labeling, and an agent backend matches queried names and returns risk information and scores.
Role:
Built the backend and AWS infrastructure
Highlights:
- VPBank Technology Hackathon 2025 finalist
- Overcame nearly 80 teams to present in the final
System for judging which of two stories is narratively closer to an anchor (abstract theme, course of action, outcome), using contrastive fine-tuning of sentence transformers trained on synthetic data. Character names are pseudonymized with consistent placeholders so the model learns narrative structure rather than names.
Role:
Co-first author (equal contribution), team ttda704; repository owner
Highlights:
- Track B: 6th place, accuracy 68.75% (top system 72.00%)
- Track A: accuracy 69.25%
- Pseudonymizing character names gave consistent gains
Track 2 (Transportation Safety Understanding and Captioning, Sim2Real) asks for detailed pedestrian and vehicle captions across five phases of each traffic-safety scenario, plus multiple-choice VQA. Models may be trained only on the synthetic SynWTS digital twin and are evaluated on real WTS videos. LoRA-fine-tuned InternVL2.5-8B with global scene frames plus local crops around pedestrian and vehicle boxes.
Role:
Adapted and extended an open-source baseline; all 20 commits after baseline import
Highlights:
- CVPR Workshop challenge under strict Sim2Real rule
- Fine-grained captions for 454 test scenarios, ~19.6k VQA questions
- Inference on a single 24 GB GPU (RTX 3090)
Classifies Vietnamese social-media posts into sarcasm in text only, in image only, in both, or none. Two approaches: ViSoBERT (text) + BEiT (image) fusion, and an end-to-end fine-tuned CLIP with a custom multimodal classifier (focal loss with class weights, mixed precision). Predictions combined by voting.
Role:
Team WEBUFF member; repository owner
Highlights:
- 10th of 43 teams on the private test (UIT DSC 2024)
- Handled low-resource, imbalanced, 4-class text+image task
A manager agent decomposes a user request, discovers worker agents (Researcher, Designer) through an on-chain registry, negotiates prices with LLM reasoning, locks funds in escrow, validates the deliverables and releases payment gaslessly. Agents mint NFT 'passports' (DIDs). A React observer dashboard streams the whole process live.
Role:
Repository owner; all commits by me
Highlights:
- Three smart contracts deployed on Kite AI Testnet
- Built for Kite AI Global Hackathon 2026
- Account abstraction (EIP-4337) and gasless EIP-3009 payments
More Projects
Additional projects in NLP, machine learning, and software development.
Experience
Professional experience in software engineering and AI development.
Platform designed to streamline core banking operations, including disbursement, lending, and debt collection, improving efficiency and reducing processing time.
- Scaled to 3,000 concurrent users (CCU) across the enterprise, ensuring high availability and system reliability.
- Reduced disbursement processing time by 40%, accelerating operational workflows and enhancing service delivery.
An automated agent that acts as a Junior Tester, capable of executing test cases and handling tasks such as writing requirements, test cases, test scripts and test reports.
- Built an end-to-end pipeline for automated test-script generation and execution using LLMs integrated via the Model Context Protocol (MCP).
- Achieved a 70% reduction in manual testing time, shortening release cycles.
Publications & Awards
Peer-reviewed research papers and competition achievements.
Publications
Proceedings of the 1st Joint Workshop on Large Language Models and Structure Modeling (XLLM 2025) @ ACL 2025, Vienna, Austria
Proceedings of the 20th International Workshop on Semantic Evaluation (SemEval-2026), San Diego, CA, USA
NTIRE Workshop @ CVPR 2025
Awards & Competitions
2nd Place — DocIE Challenge @ XLLM Workshop, ACL 2025
UIT-SHAMROCK
Finalist — AI Challenge HCMC 2026 (AIC 2026)
700+ teams
Finalist — VPBank Technology Hackathon 2025
Overcame nearly 80 teams
6th Place — SemEval-2026 Task 4 (Narrative Similarity), Track B
ttda704, accuracy 68.75%
8th Place — SemEval-2026 Task 6 (Political Evasion), Subtask 2
ttda704, Macro F1 0.5147
10th Place — UIT Data Science Challenge 2024 (Sarcasm Detection)
WEBUFF, 43 teams
Skills
Technical skills across AI/ML, software engineering, and cloud infrastructure.