About

I am a Senior Applied Scientist at Amazon working on advancing AI responsibly through evaluation, benchmark development, and automated discovery of model capabilities and weaknesses. My interests span benchmark curation, robust evaluation metrics, red-teaming and jailbreak methodologies for uncovering model vulnerabilities, and building agentic systems and LLMs for high-impact applications such as healthcare.

I completed my Ph.D. in Computing and Information Sciences at the Rochester Institute of Technology (RIT) under Dr. Linwei Wang, working at the intersection of machine learning (Bayesian modeling, optimization, generative models, graph convolutional networks) and computational healthcare — specifically personalization and uncertainty quantification for 3D cardiac electrophysiology models.

I am open to collaborations on AI safety, model evaluation, and trustworthy ML. Reach me at jwala [dot] dhamala [at] gmail [dot] com.

Research Interests

Responsible AI. I work on the responsible development of production and agentic LLM systems — from safety documentation for the Amazon Nova model family, to uncovering deceptive behaviors in long-horizon interactions, to grounding agentic reasoning in verifiable knowledge sources, to shaping the trustworthy-NLP research agenda through the TrustNLP workshop series. [Amazon Nova] [LH-Deception] [Tree-of-Traversals] [TrustNLP Retrospective]

Evaluation & Benchmarking. I design rigorous benchmarks and metrics for evaluating LLM and agentic capabilities, safety, fairness, and dialectal robustness. [BOLD] [TANGO] [Multi-VALUE] [Agentic Benchmarks]

Discovering Capabilities & Limitations. I red-team and adversarially probe models to surface emergent failures — deception auditing, joint adversarial prompting, ambiguity in text-to-image generation, and disconnects between intrinsic and extrinsic fairness metrics. [DECOR] [JAB] [T2I-Ambiguity] [Intrinsic vs Extrinsic Fairness]

AI for Healthcare. I applied ML (Bayesian optimization, graph generative models, uncertainty quantification, sequence modeling) to personalized cardiac electrophysiology and clinical decision support. [Bayesian Optimization] [Graph BO-VAE] [S2S-LSH] [GP-MCMC]

Recent News

  • July 2026: Co-organizing the 6th TrustNLP Workshop at ACL 2026 in San Diego.
  • 2026: LH-Deception on LLM deceptive behaviors in long-horizon interactions accepted at ICLR 2026.
  • Feb 2025: Serving as Area Chair for ACL Rolling Review (ARR).
  • 2025: Paper on best practices for agentic benchmarks accepted at NeurIPS 2025 Datasets and Benchmarks Track.
  • 2025: Released the Amazon Nova Technical Report.
  • May 2025: Organized TrustNLP workshop at NAACL 2025.
  • 2024: Tree-of-Traversals — paper led by our intern Elan on zero-shot reasoning with knowledge graphs accepted at ACL.

See all news for earlier updates.