LLM Reasoning & Trust · FSR Benchmark
Reliable Multi-Source Reasoning for LLMs (FSR Benchmark)
- Developed the FSR benchmark comprising 1,000 multi-source fact bundles to evaluate trust, temporal reasoning, and conflict resolution across 7B–72B LLMs using structured metadata.
- Investigated how LLMs decide what to trust when sources disagree, identifying a reliance on superficial heuristics and sycophancy over genuine epistemic reasoning.
Zeus-ML Profiling · LLaMA-3 8B
Reliability-Resource Tradeoffs in Structured Synthetic Data
- Instrumented a LLaMA-3 8B pipeline with Zeus-ML across 1,980 generation calls to profile multiobjective tradeoffs among validity, GPU energy, and latency.
- Demonstrated that schema enforcement with repair reduces GPU energy consumption by 47.9% and latency significantly.
Resource Efficiency · 1.5B–72B Models
RANGE: Resource-Aware Generation Efficiency
- Introduced an instrumented characterization framework with a usable-yield metric family (UR/J, UR/$, UR/s) applied to ten models (1.5B–72B).
- Discovered a format-dependent "CSV scale-efficiency penalty," proving that larger models consume disproportionately more energy for structured generation despite higher success rates.