Profile
Anthony K. H. Tung is a Professor at the National University of Singapore, working across database systems, data mining, information retrieval, and trustworthy, deployable AI for operational decision-making.
His current unifying theme is Prudent AI for High-Stakes, Data-Scarce Operational Systems: right-sized, interpretable, auditable, and robust AI for domains where data are scarce, labels are limited, events are rare, and deployment constraints matter.
The program threads interpretable methods into reliable systems and lightweight, deployable tools across text, time-series, tabular, and longitudinal data — connecting methods, systems, and measurable operational outcomes across manufacturing, cybersecurity, and finance/crypto. A second strand views AI as connective infrastructure that elevates collective human potential.
Research at a Glance
Research Interests & Emerging Themes
Prudent AI
Right-sized intelligence aligned to context, cost, and purpose — modular, governable "AI boxes" that lean teams can compose, audit, and deploy.
Just-in-Time AI
On-demand construction of lightweight models from scarce or streaming data for early anomaly and event detection, with controllable time-series generation for rapid adaptation.
White-Box AI
Interpretable embeddings, explainable ranking, and transparent analytics that support contestability, auditing, and trustworthy decision support.
Diversified View / RAG
Retrieval and RAG that surface complementary evidence beyond similarity — diversity-aware re-ranking and interpretable metrics to counter redundancy and echo chambers.
Trustworthy Information
Half-truth detection, omission-aware fact verification, and modeling the evolution of misinformation; robustness and safety of LLM reasoning.
Small-Data Event Models
Event-centric modeling under limited labels, rare events, and non-stationarity — combining weak supervision, synthetic data, correlation analysis, and careful evaluation.
DeepConnect
AI as shared connective infrastructure for collective sensemaking and prudent cross-domain execution — knowledge graphs, provenance-linked boundary objects, and governance.
Contributions by Pillar
Interpretable Text Embeddings
CQG-MBQA builds interpretable embeddings from contrastively generated yes/no questions; GSTransform adapts pre-computed embeddings to user instructions via guided space transformation; PRISM produces interpretable political-bias embeddings; QIME extends the idea to ontology-grounded medical text.
Trustworthy Information
DiversiNews and NEWSCOPE retrieve relevant-yet-diverse news to counter echo chambers; TRACER (with POLITIFACT-HIDDEN) formulates half-truth detection; RADAR performs omission-aware verification via role-anchored multi-agent debate; MPCG models misinformation evolution through persona-conditioned generation.
Time-Series under Scarcity
CAD detects early anomalies in sensor-based multivariate time series through correlation changes and affected-sensor localization; EADS exposes CAD through an interactive visual system; CTS supports controllable time-series generation for interpolation and extrapolation.
Benchmarks, Assistants & Workflows
TSGBench standardizes time-series-generation evaluation; CTBench tailors it to cryptocurrency series; TSGAssist uses LLMs and RAG for recommendations; TS-Agent automates financial time-series modeling via structured agentic workflows and reflective feedback.
Privacy, Security & Tabular Analytics
LDSS detects leaked tabular data used in black-box training via local distribution-shifting synthesis; RFOD performs random-forest outlier detection for mixed-type tabular data; related work extends self-explaining, weight-informed analytics to mixed-type tabular clustering.
Safety of LLM Reasoning
The Criteria Attack instantiates Reasoning Hijacking — injecting spurious decision criteria that subvert model judgments while leaving the task goal intact — exposing a blind spot in defenses that only detect goal deviation.
Deployment & Application Domains
Wafer Fabrication
- Virtual metrology & process prediction
- Excursion / drift detection & yield-risk ID
- Rare-regime anomaly detection via synthetic rare states
Identity & Access
- Malicious-login & IAM anomaly detection
- Explainable alerts under edge constraints
- Privacy-preserving augmentation & audit
Finance & Crypto
- Market & crypto simulation benchmarks
- Time-series forecasting, generation & stress-testing
- Benchmark-driven model selection; LLM workflows
Current Research Initiatives
AI Prudent Technologies (APT)
A modular ecosystem of lightweight, explainable, and governance-aware AI "boxes" that reduce compute and data demands while improving oversight and trust.
Just-In-Time.AI
Just-in-time construction of lightweight multivariate time-series models for early event detection — controllable synthetic generation for rare conditions, robust augmentation, plug-and-play edge deployment, and conversational operator interfaces.
Explainable Diversified Search in RAG
Extending RAG beyond similarity-only retrieval by balancing relevance, diversity, and user-guided coverage through hybrid retrieval, contextual re-ranking, interpretable embeddings, and visible diversity.
Half-Truth Detection & Omission-Aware Verification
Detecting presented evidence, hidden evidence, inferred claim intent, and the causal impact of missing context — including role-anchored multi-agent debate (RADAR).
Misinformation Evolution Modeling
Persona-conditioned, multi-round generation frameworks simulating how misleading claims evolve through ideological reframing, semantic drift, and audience-specific framing — enabling stress tests for detectors.
Interpretable Representation Learning
Interpretable text embeddings via automatically generated questions, instruction-guided space transformations, and bias-aware cross-encoders that expose human-readable dimensions.
Controllable Generation, Benchmarks & Assistants
Making time-series generation controllable and benchmarkable through CTS, TSGBench, CTBench, TSGAssist, and finance-oriented agentic workflows.
Controllable Unlearning & Privacy Auditing
Selective removal of data via restricted usage policies, data/model capsules, unlearning algorithms, and auditing — with detection of leaked tabular data via synthetic injection and black-box querying.
AI Mediator for Cross-Domain Collaboration (DeepConnect / Synapse)
Cross-domain mediation for experts who use different conceptual languages — using domain corpora and knowledge graphs to establish a shared framework and intervene when miscommunication arises.
Selected Publications
Full bibliography & citation counts → DBLP · Google Scholar · ACL Anthology · recent papers (combined PDF)
Patents & Applications
Software, Benchmarks & Artifacts
Awards & Recognition
Media & Public Engagement
2026
2025
2024 & Earlier
Service, Teaching & Engagement
Professional Service
- Research PC Co-Chair, VLDB 2012 · Vice PC Chair, ICDE 2012
- Workshop Chair, VLDB 2010 · Poster Chair, WWW 2010 · Vice PC Chair, ICDE 2009
- Vice PC Chair (Pre-processing), SIAM SDM 2007 · Co-PC Chair, COMAD 2006
- PC member & reviewer for SIGMOD, VLDB, ICDE, KDD, ICML, ACL/EMNLP, and IEEE TKDE
Teaching & Mentoring
- CS5344 — Big-Data Analytics Technology; graduate & undergraduate modules in database systems, data mining, search & AI systems
- Supervision across Ph.D., Master of Computing & undergraduate research — time-series, trustworthy NLP, retrieval, privacy & AI deployment
- Research leadership over teams on interpretable embeddings, diverse retrieval, fact verification, time-series generation, anomaly detection & financial AI
Industry Partners
Government Engagement
Appointments & Leadership
Education
On Research & Learning
To be constantly learning and practicing what you learn — isn't it the greatest joy of all? To have bosom friends who share your learning from all over the world — doesn't it bring the greatest happiness? To be unknown to people and yet remain undisturbed — isn't that the mark of the most virtuous? — The Analects of Confucius (论语), Chapter 1:1
The world is beautiful because there are similarities and differences. Similarities let us understand one another; differences make each of us unique and interesting.
In human relationships we respect friendship and seniority; in science and mathematics, we respect the truth.
It is more satisfying to help the weak to improve than to make the strong stronger.
The beauty of life comes from having choices among the no-choices — and from what we do with the talent we are given.
好学近乎知,力行近乎仁,知耻近乎勇。To love learning is near to wisdom; to act with vigor is near to benevolence; to know shame is near to courage.— Confucius (孔子) · The Doctrine of the Mean
君子和而不同。The noble-minded seek harmony, not uniformity.— Confucius (孔子) · Analects 13:23
言必信,行必果,硁硁然小人哉!True to every word, firm in every deed — unbending as struck stone; a lesser man, perhaps, yet worthy still.— Confucius (孔子) · Analects 13:20
尽信书,则不如无书。To trust every book completely is worse than to have no books at all.— Mencius (孟子) · Book 7B
我善养吾浩然之气。I am skilled at nurturing my vast, flowing spirit.— Mencius (孟子) · Book 2A
反者道之动,弱者道之用。Reversal is the movement of the Dao; yielding is the working of the Dao.— Laozi (老子) · Tao Te Ching 40
知人者智,自知者明。To know others is wisdom; to know oneself is enlightenment.— Laozi (老子) · Tao Te Ching 33
物物而不物于物。Master things — and do not be mastered by them.— Zhuangzi (庄子) · The Mountain Tree
无用之用,方为大用。The use of the useless is the greatest use of all.— Zhuangzi (庄子) · The Human World