Completion Date
Spring 5-4-2026
Document Type
Thesis
Degree Name
Doctor of Philosophy (PhD)
Program or Discipline Name
Data Sciences
First Advisor
Kayden Jordan, Ph.D.
Second Advisor
Srikar Bellur, Ph.D.
Third Advisor
Sangwhan Cha, Ph.D.
Abstract
This dissertation focuses on three applied challenges in machine learning and artificial intelligence: predicting social media content virality, evaluating safety in tool using large language model agents, and standardizing model validation in regulated settings. Across these domains, the dissertation argues that evaluation in applied AI must be matched to context, combining predictive performance with the relevant evidence needed for interpretability, safety, reproducibility, and governance readiness. First, Decoding Reddit Memes Virality extracts and analyzes 16,968 posts using computer vision, natural language processing, and gradient boosting to identify visual, textual, and temporal signals associated with virality. Second, The Verifier Tax designs and evaluates baseline, planning-integrated, and policy mediated agent architectures to quantify safety-performance tradeoffs, showing that runtime safety enforcement can block unsafe actions but often fails to recover safe task completion, revealing a persistent Safety-Capability Gap. Third, this dissertation develops TanML, an open source toolkit that integrates data profiling, data preprocessing, feature power ranking, model development, model evaluation, explainability, drift analysis, and audit ready report generation for tabular machine learning models. Together, these contributions provide empirical evidence, reusable datasets, and practical software for applied data science. They are supported by publicly available code and data on GitHub, including a dataset of 16,968 Reddit memes on the Open Science Framework. The TanML toolkit is additionally released on PyPI. The findings clarify how multimodal features shape content outcomes, where current agent safety mechanisms break down, and how validation workflows can be made more consistent and governance ready in documentation heavy environments.
Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License
Recommended Citation
Sah, T. (2026). Evaluation in Applied AI: Predictive Modeling, Agent Safety, and Model Validation. Retrieved from https://digitalcommons.harrisburgu.edu/dandt/103