Research
Agent evaluation, selective risk, and multimodal neural decoding.
Three 2026 preprints and earlier journal work. Two of the preprints ship public, reproducible code; where a number here came from a benchmark, the benchmark is linked.
The record in four numbers
- 3
- 2026 preprints — two shipping public, reproducible code
- 14,750
- agent execution traces in AgentProp-Bench
- 13,859
- document fields audited for selective risk
- 84.8%
- fused EEG–fMRI decoding accuracy, within-subject
Preprints — 2026
The two arXiv papers release the harness that produced their numbers, so the figures are reproducible rather than reported. The third is under review as the benchmark companion to the first.
- 2026
13,859 fields from 800 CORD receipts; seed-pinned, regression-gated harness.
↗ arXiv:2608.14639 [cs.LG, cs.AI, cs.CL]↗ verifydoc (Apache-2.0)case study →
- 2026
AgentProp-Bench: 14,750 execution traces from 13 LLM agents across 4 domains.
↗ arXiv:2604.16706 [cs.AI, cs.CL, cs.MA]↗ agenthallu-benchcase study →
- 2026
Gurram, B. VerifyDocBench: Measuring Per-Field Calibration, Selective Risk, and Grounding for Document Extraction — at Scale, Across Models and Languages.
Companion benchmark paper to arXiv:2608.14639.
Under review
The complementarity claim
EEG knows when. fMRI knows where. Fused, they beat either alone.
M.S. thesis
A Multimodal Neuroimaging Method for the Prediction of Visual Stimuli
University of Cincinnati · M.S. Computer Science · GPA 3.95/4.0 · Advisor: Prof. Vikram Ravindra
EEG–fMRI fusion framework using temporal convolutional networks for neural decoding. 70-channel EEG + 3T fMRI (Wakeman–Henson dataset); SSS, ICA/wavelet-ICA, and DWT preprocessing; MNI-normalized fusiform/occipital ROIs. The combined model reached 84.8% within-subject and 81.1% cross-subject (LOSO) accuracy with ROC-AUC 0.93/0.90 — against 65.5% EEG-only and 74.6% fMRI-only, beating GRU, LSTM-CNN, and SVM baselines.
Reproducible artifacts
Both 2026 arXiv preprints ship their harness. verifydoc is seed-pinned and regression-gated under Apache-2.0; agenthallu-bench publishes its code and data.
Research artifact · arXiv:2608.14639
verifydoc
arXiv preprint · 14 pages
Reference implementation for “Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays.” It shows that the natural per-field accept/review procedure silently violates its own selective-risk guarantee on real documents.
- 13,859
- fields evaluated
- 800
- CORD receipts
- 3
- failure modes diagnosed
- 49.0%
- base field accuracy
Research artifact · arXiv:2604.16706
agenthallu-bench
arXiv preprint · 11 pages, 4 figures, 8 tables
Code and data for “Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents.” It audits the widely-assumed reliability of automated evaluation for tool-using LLM agents against human annotation.
- 14,750
- execution traces
- 13
- LLM agents audited
- 4
- task domains
- 9 + 4
- proprietary + open-weight models
Journal papers — 2021
Earlier co-authored work, from the undergraduate years, listed for completeness. No public identifier is quoted for these beyond the volume and page range the record holds.
- 2021
Gurram, B. A Review on Secure Data Transmission for Banking Application using Machine Learning. With Reddy, M.D. & Thatikonda, M.
IJEAT, 10(5), 182–186
- 2021
Gurram, B. Analysis of Big Data Challenges and Different Analytical Methods. With Reddy, M.D.
IJERAT, 7(3), 33–38
- 2021
Gurram, B. Challenges and Solutions for Improving SSD Performance. Et al.
IJRES, 9(8), 37–39
Keep reading
The engineering behind these →Experience & teaching →↗ Google Scholar