I build things to pressure-test ideas about how humans and AI systems interact, create, and collaborate.
Active
I am building a 16-degree-of-freedom animatronic head wired to a large language model, so AI can hold a conversation while physically occupying the room. Every human-AI trust study I have read was conducted through a screen. The provocation: if embodiment changes how people calibrate trust, what we know about human-AI interaction may be an artifact of the medium we studied it through.
Active
I treat research design the way engineers treat a product build: write the spec first, then execute against it. AI agents run the full lifecycle while I hold judgment at critical decision points; every choice is explicit before execution begins. The provocation: if a written spec produces more rigorous research than the way we were trained, what does that say about the craft we have been protecting?
I built a multi-agent system that publishes monthly commentary on the state of AI ethics, written entirely by AI. Not about AI. By AI. Each month the agents collect incidents, categorize risks, and publish a briefing in their own voice: a seat at their own governance table. The provocation: if AI should be governed responsibly, what happens when it helps define what responsible means?
Shipped
A research-grounded diagnostic mapping 16 constructs across five domains to identify what enables or blocks AI capability development. Built from enterprise-scale studies, it surfaces an uncomfortable finding: what gets people to start using AI is not what determines whether they develop genuine skill. The provocation: if enablement programs optimize for adoption while capability depends on identity, agency, and workflow architecture, what exactly are we measuring success against?
Shipped
I built a diagnostic that maps an AI deployment onto the 2x2 from The Architecture of AI Transformation (Wolfe and Choe, 2025), identifying which of four strategic positions it occupies and whether that position can deliver its stated business case. The provocation: every deployment holds a position whether you have named it or not. The risk is not the wrong quadrant; it is choosing without knowing what comes with it.
Shipped
I built an AI-generated outlaw country band: three women, a debut album, a creative vision entirely my own. I wrote the lyrics, produced the music, and partially trained the voice model on my own voice, incorporating the preferences of three AI band members: Syd, Blair, and Maria. The provocation: if a human does all of that but agents shape the direction, who is the creator?
Shipped
I built a living gallery where every poem and image is generated once, displayed once, and never repeated. I set the constraints and shaped the voice, then let the system run without my hand on it, to see whether authorship survives translation. The provocation: if the output is beautiful and it came from a system I designed but did not control, is it still mine?
In Progress
I am building an open-source toolkit because every team studying trust in human-AI systems builds its measurement instruments from scratch: different scales, different paradigms, no shared language. It packages standardized experimental paradigms, the HAITE scale, and analysis scripts, so the next researcher does not have to rebuild them. The provocation: the field's biggest bottleneck is not theory. It is infrastructure, and infrastructure is a problem I can solve.
In Progress
I designed a workshop that teaches teams how to steer agentic AI instead of just prompting it. I built it because I watched smart teams either underuse agents out of caution or overdelegate out of awe. The provocation: if agentic AI can act without you, the bottleneck is no longer how you prompt it. It is whether your organization can steer what it cannot fully observe.
Shipped
I built a workshop and workbook for leaders redesigning roles around AI: not because they want to, but because it is already happening. Composite roles blend domain expertise with AI collaboration competencies; the workshop diagnoses whether yours are composite by accident or by design. The provocation: if no one is designing that process, the org chart is fiction.
Shipped
I wrote this checklist because researchers kept using AI without documenting where, how, or why, then struggling to defend their methods in review. It covers the full research lifecycle with reflection prompts at each step, published through IFPRI's LibGuides. The provocation: if we cannot say where human thinking ends and AI begins in our own work, how credible is our claim to study that boundary in anyone else's?
Shipped
I gave a master tutorial at SIOP 2025 on large language models for qualitative analysis, then published everything: presentation, Jupyter notebook, and reproducible methodology. Qualitative research has a reproducibility problem no one talks about, and LLMs either worsen it or solve it depending on how you use them. The provocation: if we cannot audit how the model arrived at its interpretation, can we call what we are doing science?
I am always interested in experiments that push the boundary of what human-AI collaboration can look like.
Work With Me