ĐTÂ

Đinh Thiên Ân

Software Engineer at VPBank · AI Engineer — LLMs, RAG & Agents, Vision-Language Models

Ho Chi Minh City, Vietnam

Software Engineer at VPBank in Ho Chi Minh City and Computer Science graduate (B.Sc., GPA 3.5/4.0) of the University of Information Technology (UIT – VNU-HCM). Previously built an LLM + MCP test-automation agent as an AI Engineer Intern at TMA Solutions. Co-author of ZeroSemble (XLLM Workshop @ ACL 2025; 2nd place, DocIE shared task) and of two SemEval-2026 system papers (6th place in SemEval-2026 Task 4, Track B). Builds RAG, multi-agent, video-retrieval and vision-language systems; his team reached the final of the AI Challenge HCMC 2026 among 700+ teams.

Featured Projects

Research projects, competition entries, and production systems spanning NLP, computer vision, multi-agent systems, and full-stack development.

ZeroSemble — Zero-Shot Document Information Extraction with LLM Ensembles
NLP Research / Information Extraction · 2025

Zero-shot, two-stage pipeline for document-level entity and relation extraction. First, entities are extracted by three heterogeneous LLMs (DeepSeek-R1 distill, Llama-3.3-70B, Qwen-2.5-32B) and consolidated through deduplication and majority-vote typing. Then relations are extracted with prompts that constrain head and tail to the consolidated entity set. No domain-specific training is required.

Role:

Co-author (2nd author); repository owner

Highlights:

  • 2nd place, DocIE shared task @ XLLM Workshop, ACL 2025 (team UIT-SHAMROCK)
  • Published in ACL Anthology (2025.xllm-1.25)
  • F1 scores: EI 55.65, EC 26.11, REG 4.19, RES 4.01, Overall 22.49
PythonLLM APIs (Groq)DeepSeek-R1Llama-3.3-70BQwen-2.5-32BPrompt Engineering
AIC Mentos — Interactive Video Event Retrieval
Multimodal Retrieval / Full-Stack System · 2026

Interactive search system for finding events in a large video archive. The index covers 1,487 videos in 125 groups and 524,891 keyframes (news, CCTV traffic cameras, cycling). Features: Vietnamese text-to-image visual search on SigLIP2 with LLM query translation, OCR text search, ASR transcript search, object-detection filtering, temporal multi-event search, voice queries, and a built-in DRES submission bar.

Role:

Built the entire system

Highlights:

  • Finalist, AIC 2026 (AI Challenge HCMC), 700+ teams
  • Index: 1,487 videos · 125 groups · 524,891 keyframes
  • Latency: OCR ~0.17s, ASR ~0.11s, object detection <1s, temporal search ~4s
SigLIP2LLM Query TranslationOCR SearchASR SearchObject DetectionReact+1 more
Intelligent Risk Analyzer for AML Review
Agentic AI / FinTech / Cloud · 2025

Searches, gathers and analyzes online information (news, social media, public data) about customers from their personal or business identification details. It builds a risk blacklist from crawled news with LLM-based extraction and risk labeling, and an agent backend matches queried names and returns risk information and scores.

Role:

Built the backend and AWS infrastructure

Highlights:

  • VPBank Technology Hackathon 2025 finalist
  • Overcame nearly 80 teams to present in the final
LangChainAWS BedrockAWS SageMakerAWS LambdaAWS S3/EC2DynamoDB+2 more
SemEval-2026 Task 4 — Narrative Similarity
Representation Learning / Contrastive Learning · 2026

System for judging which of two stories is narratively closer to an anchor (abstract theme, course of action, outcome), using contrastive fine-tuning of sentence transformers trained on synthetic data. Character names are pseudonymized with consistent placeholders so the model learns narrative structure rather than names.

Role:

Co-first author (equal contribution), team ttda704; repository owner

Highlights:

  • Track B: 6th place, accuracy 68.75% (top system 72.00%)
  • Track A: accuracy 69.25%
  • Pseudonymizing character names gave consistent gains
PythonPyTorchSentence-TransformersOpenAI APIContrastive LearningWeights & Biases
AI City Challenge 2026 Track 2 — Sim2Real Traffic-Safety Video Captioning & VQA
Computer Vision / Vision-Language Models · 2026

Track 2 (Transportation Safety Understanding and Captioning, Sim2Real) asks for detailed pedestrian and vehicle captions across five phases of each traffic-safety scenario, plus multiple-choice VQA. Models may be trained only on the synthetic SynWTS digital twin and are evaluated on real WTS videos. LoRA-fine-tuned InternVL2.5-8B with global scene frames plus local crops around pedestrian and vehicle boxes.

Role:

Adapted and extended an open-source baseline; all 20 commits after baseline import

Highlights:

  • CVPR Workshop challenge under strict Sim2Real rule
  • Fine-grained captions for 454 test scenarios, ~19.6k VQA questions
  • Inference on a single 24 GB GPU (RTX 3090)
PythonPyTorchInternVL2.5-8BLoRA / PEFTDeepSpeed ZeRO-3Hugging Face+1 more
Multimodal Sarcasm Detection for Vietnamese Social Media
Multimodal Deep Learning / Competition · 2024

Classifies Vietnamese social-media posts into sarcasm in text only, in image only, in both, or none. Two approaches: ViSoBERT (text) + BEiT (image) fusion, and an end-to-end fine-tuned CLIP with a custom multimodal classifier (focal loss with class weights, mixed precision). Predictions combined by voting.

Role:

Team WEBUFF member; repository owner

Highlights:

  • 10th of 43 teams on the private test (UIT DSC 2024)
  • Handled low-resource, imbalanced, 4-class text+image task
PythonPyTorchViSoBERTBEiTCLIPHugging Face+1 more
AgentMesh — Autonomous Multi-Agent Marketplace on Kite AI
Multi-Agent Systems / Web3 / Full-Stack · 2026

A manager agent decomposes a user request, discovers worker agents (Researcher, Designer) through an on-chain registry, negotiates prices with LLM reasoning, locks funds in escrow, validates the deliverables and releases payment gaslessly. Agents mint NFT 'passports' (DIDs). A React observer dashboard streams the whole process live.

Role:

Repository owner; all commits by me

Highlights:

  • Three smart contracts deployed on Kite AI Testnet
  • Built for Kite AI Global Hackathon 2026
  • Account abstraction (EIP-4337) and gasless EIP-3009 payments
React 19ViteTailwind CSSFastAPIWeb3.pyGoogle Gemini+3 more

More Projects

Additional projects in NLP, machine learning, and software development.

Legal Graph RAG — Vietnamese Legal Retrieval
Knowledge Graphs / Information Retrieval · 2026

Graph RAG for Vietnamese legal texts with validity checking via REFERENCES/EXCEPTS and REPLACES/AMENDS edges. Hybrid dense + BM25/SPLADE retrieval with RRF fusion.

PythonNeo4jDense EmbeddingsBM25/SPLADE

Ships with ablation and error-analysis scripts

CartMate — RAG Product-Search Chatbot
RAG / LLM Application · 2024

RAG chatbot for product search, FAQs, order tracking and personalized recommendations using hybrid search (BM25 + semantic).

RAGFastAPIReact/Node.jsMongoDB

Context-aware product searches and recommendations

Building Energy Efficiency Prediction
Classical ML / Full-Stack · 2025

Predicts heating and cooling loads from architectural parameters. Flask API with Streamlit and React UIs.

PythonScikit-learnXGBoostFlask

Random Forest R² > 0.99 for both heating and cooling load

Emotion Recognition for Vietnamese Social Media
NLP (Vietnamese) · 2024

Fine-tuned ViSoBERT on the UIT-VSMEC corpus (6,927 sentences, 7 emotions).

PythonViSoBERTPyTorchStreamlit

Accuracy 66%, weighted F1 66%, macro F1 64%

SemEval-2026 Task 6 — Political Evasion Detection
LLM Reasoning / NLP · 2026

Classifies clarity and evasion strategies in question-answer pairs using structured chain-of-thought prompting.

Qwen3DeepSeek-V3.2Grok-4-FastQLoRA

Rank 8/33 on Subtask 2 (Macro F1 0.5147)

Experience

Professional experience in software engineering and AI development.

Software Engineer
VPBank
Sep 2025 – PresentHo Chi Minh City, Vietnam
Smart Tools Suite

Platform designed to streamline core banking operations, including disbursement, lending, and debt collection, improving efficiency and reducing processing time.

  • Scaled to 3,000 concurrent users (CCU) across the enterprise, ensuring high availability and system reliability.
  • Reduced disbursement processing time by 40%, accelerating operational workflows and enhancing service delivery.
AI Engineer Intern
TMA Solutions
Mar 2025 – Jul 2025Ho Chi Minh City, Vietnam
Testing Copilot

An automated agent that acts as a Junior Tester, capable of executing test cases and handling tasks such as writing requirements, test cases, test scripts and test reports.

  • Built an end-to-end pipeline for automated test-script generation and execution using LLMs integrated via the Model Context Protocol (MCP).
  • Achieved a 70% reduction in manual testing time, shortening release cycles.

Publications & Awards

Peer-reviewed research papers and competition achievements.

Publications

DocIE@XLLM25: ZeroSemble — Robust and Efficient Zero-Shot Document Information Extraction with Heterogeneous Large Language Model Ensembles
Nguyen Pham Hoang Le, An Dinh Thien, Son T. Luu, Kiet Van Nguyen

Proceedings of the 1st Joint Workshop on Large Language Models and Structure Modeling (XLLM 2025) @ ACL 2025, Vienna, Austria

2025View
ttda704 at SemEval-2026 Task 4: Modeling Narrative Structures via Pseudonymization and Multi-View Sentence Alignment
Tai Tran Tan*, An Dinh Thien* (*equal contribution)

Proceedings of the 20th International Workshop on Semantic Evaluation (SemEval-2026), San Diego, CA, USA

2026View
ttda704 at SemEval-2026 Task 6: Structured Chain-of-Thought Prompting for Political Evasion Detection
Tai Tran Tan, An Dinh

Proceedings of the 20th International Workshop on Semantic Evaluation (SemEval-2026), ACL, San Diego, July 2026, pages 2817–2829

NTIRE 2025 Challenge on Day and Night Raindrop Removal for Dual-Focused Images: Methods and Results (co-author, team method)
Multiple authors including An Dinh Thien

NTIRE Workshop @ CVPR 2025

2025View

Awards & Competitions

2nd Place — DocIE Challenge @ XLLM Workshop, ACL 2025

UIT-SHAMROCK

2025

Finalist — AI Challenge HCMC 2026 (AIC 2026)

700+ teams

2026

Finalist — VPBank Technology Hackathon 2025

Overcame nearly 80 teams

2025

6th Place — SemEval-2026 Task 4 (Narrative Similarity), Track B

ttda704, accuracy 68.75%

2026

8th Place — SemEval-2026 Task 6 (Political Evasion), Subtask 2

ttda704, Macro F1 0.5147

2026

10th Place — UIT Data Science Challenge 2024 (Sarcasm Detection)

WEBUFF, 43 teams

2024

Skills

Technical skills across AI/ML, software engineering, and cloud infrastructure.

Languages
PythonC/C++JavaScript (Node.js)SQL
AI / Machine Learning
NLPRAG PipelinesOCRComputer VisionMultimodal LearningLLM Optimization (quantization, GGUF, llama.cpp)
Frameworks & Libraries
PyTorchTensorFlowLangChainFastAPIHugging Face TransformersScikit-learnSentence Transformers
Backend & Infrastructure
PostgreSQL (pgvector)RedisRabbitMQMongoDBDockerGitLinux/Ubuntu
Cloud Platforms
AWS (Bedrock, SageMaker, Lambda, EC2, S3)Google Cloud Platform
Tools & Protocols
Model Context Protocol (MCP)Vector DatabasesMulti-stage Docker Builds

Education & Certifications

Education

B.Sc. in Computer Science
University of Information Technology (UIT), VNU-HCM
2022 – 2025Ho Chi Minh City, Vietnam
GPA: 3.5/4.0Graduated

Certifications

Applications of AI for Anomaly Detection

NVIDIA Deep Learning Institute

TOEIC Listening & Reading 790/990

ETS