James M. Bradley Jr., Ph.D.
Staff UX Researcher. I run rigorous mixed-methods research on product timelines, and I build the systems that let insight scale.
Research repositories, measurement instruments, and governed AI-augmented workflows for a top-5 US sportsbook. Twelve years of research across product UX, applied behavioral science, and institutional analytics.
What I build around the research

Infrastructure
A cross-team research repository of 2,100+ documents across 32 product areas, merged under one provenance model so findings stay attributable and queryable. Standardized, KPI-linked measurement scales reused across product teams.
AI, governed
A self-serve Slack research agent that answers with cited verbatims, evidence dates, and honest confidence flags. An intake agent that scores product concepts against customer evidence for leadership prioritization. An opportunity-backlog role that tracks needs surfacing outside any one study's scope, and a prioritization role that prepares the standing research-and-strategy forum with product leadership. Seven automated workflows keeping the corpus current.
Craft
The full mixed-methods range, matched to the decision at hand: controlled experiments when a design choice needs a defensible winner, large-n surveys with real data-quality gates when the roadmap needs an order, in-depth and in-person work when the question is why. A standing research program keeps the evidence current, and a Ph.D. in social psychology is what tells me which method a question deserves and what threatens its validity.
Case studies
Six studies chosen to show range. Every number is one I can walk you through; employer-specific figures are generalized on purpose.
Building an AI-Augmented Research Practice
A scattered archive became a governed, self-serve knowledge system: 2,100+ documents, five AI research roles, and seven scheduled workflows the whole team uses.
Read the case study →The Long Game
Two years of cross-study signal consolidated into a buildable case; a first scoped step is now moving with leadership aligned.
Read the case study →Research at the Speed of Shipping
A standing research program that turns a product team's question into a customer-evidence readout inside two weeks: four rounds and 40+ sessions in its first three months.
Read the case study →Designing to Prevent Costly Errors
Identified the one layout that performed on both tasks and structurally prevented accidental spends of rewards currency, from 62 participants in two weeks.
Read the case study →Knowing When to Pivot vs. Optimize
A validation request became a discovery study whose answer redirected a committed engineering investment away from optimizing a feature customers had learned to avoid.
Read the case study →Field Research Under Real Stakes
Live-decision performance and stability moved to the top of the next quarter's roadmap across three product teams, with dedicated engineering allocation.
Read the case study →Always open to a conversation about research.
Questions about any of this work, the methods behind it, or the AI-augmented practice: I'm happy to talk.
Get in touchCase studies
Research that shows the work
Research infrastructure and AI-augmented practice, evidence-based advocacy over two years, a standing research program, experimental usability craft, strategic reframing, and field research. Every number is one I can walk you through.
Building an AI-Augmented Research Practice
A scattered archive became a governed, self-serve knowledge system: 2,100+ documents, five AI research roles, and seven scheduled workflows the whole team uses.
Read the case study →The Long Game
Two years of cross-study signal consolidated into a buildable case; a first scoped step is now moving with leadership aligned.
Read the case study →Research at the Speed of Shipping
A standing research program that turns a product team's question into a customer-evidence readout inside two weeks: four rounds and 40+ sessions in its first three months.
Read the case study →Designing to Prevent Costly Errors
Identified the one layout that performed on both tasks and structurally prevented accidental spends of rewards currency, from 62 participants in two weeks.
Read the case study →Knowing When to Pivot vs. Optimize
A validation request became a discovery study whose answer redirected a committed engineering investment away from optimizing a feature customers had learned to avoid.
Read the case study →Field Research Under Real Stakes
Live-decision performance and stability moved to the top of the next quarter's roadmap across three product teams, with dedicated engineering allocation.
Read the case study →Research infrastructure · 2025–2026
Building an AI-Augmented Research Practice: From Scattered Decks to a Governed Knowledge Base the Whole Team Can Query
A governed research repository, five AI research roles including a self-serve Slack agent and an intake agent for leadership prioritization, and an automated ops pipeline, built without lowering the evidence bar
- Role
- Sole UX researcher for the sportsbook; designed and built the repository, standards, and every agent and workflow
- Timeline
- 2025 – present
- Methods
- Repository architecture · taxonomy design · AI agent and workflow design · governance and accuracy auditing

The problem
A fast-shipping product had years of research spread across decks, drives, and two separate research teams. Findings were hard to locate, easy to duplicate, and impossible to query across. Leadership increasingly wanted fast, evidence-grounded answers to "what do we already know about X?" and "is this concept worth building?", faster than any single researcher could assemble by hand.
The interesting problem was never "use AI." It was keeping research rigor intact while AI compresses the work: how does a fast answer still cite its sources, show its freshness, and admit its gaps?
What I built
A governed research repository. 2,100+ documents across 32 product areas and 130+ study-level workstreams in a single, host-agnostic knowledge base (portable Markdown and CSV, not locked to any tool). A provenance model (source, product line, researcher) keeps two teams' findings attributable and comparable rather than blurred together. A versioned tagging taxonomy, evidence-recency ("freshness") standards, and a PII handling and quarantine pipeline keep the shared layer clean.
Self-serve answers where the team already works. A Slack agent answers research questions from the repository during business hours. Every answer carries what a good researcher would insist on: real customer verbatims with attribution, the date and freshness of the evidence behind each finding, links to the canonical deliverable, and an explicit low-confidence flag (with escalation to me) when a decision-shaped question has thin evidence. The agent is allowed to say "the evidence can't answer this safely."
Governed AI-augmented synthesis workflows. The repeatable stages of qualitative research (session analysis, cross-participant synthesis, stakeholder communication) codified into reusable AI workflows. Not speed for its own sake: speed with a fixed quality bar. Attributed verbatims behind each finding, evidence dated per section, an explicit "where the evidence is thin" section. The practical effect is that the time from the end of data collection to a finished deliverable dropped from weeks to three or four days.
An opportunity-backlog role. Customer needs that surface outside a study's scope don't die in the margins. They are extracted into a governed register that tracks each theme across studies (first said, last said, how often, by whom) and deduplicates against the roadmap, topped up on a monthly cycle. It is how a need raised in a navigation study in 2024 and a churn study in 2025 becomes one accumulating case instead of two forgotten footnotes.
An intake agent for prioritization. Scores an incoming product concept against the combined customer-evidence corpus: an evidence-scoped priority and a separate confidence rating (so a weakly evidenced idea cannot masquerade as a strong one), every claim cited with a verbatim where one exists, conflicts surfaced rather than averaged away, and a mandatory accuracy gate. Designed to argue both for and against a concept. Researcher in the loop. Ten concept assessments delivered to date, including a cross-batch comparison across a six-concept intake.
A research and strategy prioritization role. The sportsbook research and strategy prioritization forum I co-run with product leadership every two weeks has its own agent, with its own rule set: it assembles the period's completed research into the forum's materials so the meeting starts from evidence rather than from whoever had time to prepare the deck. It is read-only against the repository, and meeting material is deliberately never ingested back into the evidence base. Each of the five roles is a distinct set of rules and judgment over the same repository; none of them is a general-purpose chatbot.
An automated operations layer. Seven scheduled workflows keep the system running without manual effort: two recurring voice-of-customer ingests, a weekly cross-team research-portal scan, a weekly auto-ingest of my own completed work with a hold-on-doubt gate (anything ambiguous, PII-flagged, or unaudited is held for human review, not filed), a weekly guarded backup-parity check that blocks suspicious deletions rather than propagating them, and the two schedules on which the Slack agent (hourly, business hours) and the forum-prep role (weekly) run. Two of the seven schedules, in other words, are how two of the five roles operate.
Responsible AI, as governance not garnish
These outputs inform real investment decisions, so the guardrails are built in: accuracy verification against source, provenance on every claim, down-weighting of stale evidence, and a responsible-gaming check on anything customer-facing. The canonical failure designed against is the confident-but-wrong statistic (a "93%" that is really 8 of 12); the accuracy gate exists to catch exactly that before it reaches a stakeholder. A standing provenance rule prevents the repository's own outputs from re-entering as evidence, so a finding can never look better supported than it is.
Q: What do we know about why users abandon live same-game parlays?
Finding. Market suspensions during live play teach users to stop trying; the pattern is learned helplessness rather than a usability gap. evidence: Jul 2025 · 12 sessions
"I don't even bother at this point." P07 · live bettor · attributed
Impact
- Turned a scattered archive into a self-serve capability the team queries directly in Slack, rather than re-running studies that already exist.
- Gave leadership a faster, evidence-grounded input to concept prioritization, with the honesty of a confidence rating attached.
- Established the standards (provenance, recency, verbatim-backing, accuracy gate) that let the team trust AI-assisted answers.
What I'd tell another researcher
This is a governance and standards problem as much as a tooling one, and it is where research leadership will increasingly live. The agents are the visible part; the taxonomy, the freshness model, the PII pipeline, and the accuracy gate are what make them safe to put in front of people who will act on the answer.
Evidence-based advocacy · 2024–2026
The Long Game: Building the Evidence Case for Personalization
Two years of signal from studies designed to answer other questions, consolidated into a buildable case, validated quantitatively, and carried until it moved
- Role
- Sole researcher throughout: ran the targeted studies, maintained the register, wrote the synthesis and design brief, ran the validation survey
- Timeline
- Aug 2024 – present
- Methods
- Unmoderated IDIs · pre/post concept survey (n = 379) · cross-project audits · research synthesis · validation survey (n = 527)

The problem
Sports betting products are built around the bet: the market, the odds, the slip. Personalization, in the sense of an app that knows which teams, leagues, and markets you care about and organizes itself around them, sits outside that frame. It kept losing prioritization decisions to nearer-term improvements. Reasonably, at first: my own survey ranked it third. But the signal kept arriving in studies designed to answer other questions, and the case eventually became one no roadmap could ignore.
This is a case study about what a researcher does with a need that keeps showing up when nobody asked about it.
The first read: honest, and third
The signal before the study. The first mentions were unprompted. A study of a fast-paced betting feature in August 2024 produced a recommendation for a few personalized markets on the page. A study of league pages that October produced one for personalized quick access in navigation. Neither study was about personalization. Both were filed, and both were tracked.
The targeted read. In late 2024 I ran personalization's first dedicated program: twelve unmoderated in-depth discovery interviews across customer segments, then a validation survey of 379 current customers across three engagement tiers, ranking six candidate features. Participants ranked before and after seeing design concepts, because stated preference alone misses latent demand: people cannot want what they cannot picture.
What I reported. Personalization ranked third of six. I said so, and I recommended cashout and data-visualization improvements first, because customers ranked them higher. But the pre/post read said something the rank alone did not: personalization was the only feature that rose significantly after exposure (mean rank 3.66 to 3.41, p < .01), and once people had seen it, 78.1% said they would use it, 63.9% expected it to raise their satisfaction, and 57.3% expected it to increase their share of wallet. The demand was real; it was latent. So the recommendation was two-part: not now, and do the qualitative design research before building, because this one will be right or wrong on the details.
The signal keeps arriving
Over the following year, personalization surfaced in study after study, most of them designed to answer something else. A live-betting customization theme recurred across three studies between late 2024 and mid-2025. A study of a narrative betting concept in August 2025 produced three separate personalization recommendations. At the October 2025 in-person research event, 26 participants raised favoriting across markets, teams, players, and leagues, a following view of starred games, and hiding sports they never bet, each theme recurring across three studies. The churn work that winter surfaced it again among high-engagement customers who had left.
None of this would have added up to anything on its own. What made it add up was the practice built around the research: every off-plan need extracted into a governed register that tracks each theme across studies, with first-said and last-said dates, instance counts, and a deduplication pass against the roadmap. The register is what let a navigation finding from 2024 and a churn finding from 2025 become one accumulating case rather than two forgotten footnotes.
The consolidation
In March 2026 I stopped waiting for the signal to be commissioned and wrote the case. Four cross-project audits (the in-person event, the lifetime-value work, a data-visualization study, and the high-engagement churn work) fed a research synthesis whose opening argument was the register's: over eighteen months, 17+ projects and 500+ participants, spanning surveys, moderated and unmoderated sessions, field observation, and design evaluation, had independently surfaced the same need, and the vast majority were not designed to study it.
The synthesis came with the thing advocacy usually lacks: a buildable shape. A three-phase roadmap: first a manually curated favorites view built around the entities customers actually follow, then contextual surfacing and behavioral learning, then promotional intelligence. And a design brief for phase one. The argument stopped being "customers want personalization" and became "here is the first thing to build, and here is why it is the right first thing."
Validation
In May 2026 a survey of 527 customers tested the first-phase scope directly. The first-phase concept was the clear preference, chosen at roughly twice the rate of the next option. Three months later, when a standing research round set a launch requirement for a live-content concept (curated, not algorithm-only), it was this survey, not the round's own nine sessions, that settled it.
Movement
A first, scoped step is now moving, with leadership aligned on it. It is not the full first phase, and it is not the hub the evidence describes. It is a start, and it is the first time the roadmap has moved on this in two years. The signal is still arriving: the most recent standing research round surfaced home-surface personalization again, unprompted, in a round designed to evaluate something else.
What I'd tell another researcher
Advocacy is not volume. It is provenance. This case survived every re-litigation because every claim in it traces to a study, a date, and a verbatim, and because the researcher making it had reported the inconvenient result first. Ranking third, and saying so, is what made "it keeps coming back" credible when it did. The other lesson is about infrastructure: a need that surfaces outside a study's scope is only evidence if something catches it. Build the register before you need the case.
Program design · 2026
Research at the Speed of Shipping: A Standing Evidence Pipeline
A continuously running moderated research program on a three-week beat, with a fixed runbook that turns a product team's Monday question into a customer-evidence readout inside two weeks
- Role
- Designed the program, wrote the playbook and runbook, and moderate and analyze every round; product and design owners brief concepts in
- Timeline
- Jul 2026 – present
- Methods
- Moderated 45-minute sessions · post-launch evaluation · discovery · concept review · returner re-tests · same-day analysis

The problem
Most research is commissioned one study at a time, and it thins out exactly where products need it most. Teams get research before they build and then nothing after launch, so the question "did the thing we shipped actually work for customers?" goes unanswered or gets answered by analytics alone. At the same time, small asks from product and design ("can you just check one thing with users?") either derailed planned work or died in a queue. Neither problem is solved by running more studies. Both are solved by changing the shape of the research: from a series of commissions to a standing program.
The design
One program, one calendar commitment. A moderated round with the product's highest-value customer segment roughly every three weeks. Three weeks is the floor, not a compromise; it is the shortest cadence at which a round can be recruited, run, analyzed, and read out without cutting a corner. The round's focus rotates with product need: post-launch evaluation of recently shipped surfaces, discovery in natural use, or concept and design review of in-flight work from several product teams. Round size flexes with the question too: the fourth round was deliberately extended to fifteen sessions to evaluate a redesigned surface across participants whose exposure to it varied by weeks. Alternating focus inside one program was deliberate, because one calendar commitment is harder to erode than two that each look optional.
A runbook that fixes the timeline. Monday of the prior week, an open intake call to product and design for research questions, first come, first served. Friday, intake closes and qualifying customers are pulled from the data warehouse. The following Monday, recruiting goes out, self-serve booking opens, and the discussion guide is finalized by slotting the round's asks into standing guide skeletons. Tuesday and Wednesday, sessions run and each is analyzed the same day. Thursday and Friday, cross-participant synthesis, readout, Slack shareout, and filing into the research repository and the opportunity backlog. Ask to answer in under two weeks.
An intake rule for the small asks. An ad-hoc question is absorbed into the next round if it fits the participant profile and needs no net-new stimuli; otherwise it waits a round or becomes its own project. The queue that used to swallow small asks no longer exists, because the next round is never more than three weeks away.
Standing infrastructure. A program playbook, round-agnostic discussion-guide skeletons for each round type, standing analysis templates, a warehouse-query recruiting pipeline with self-serve scheduling, a locked incentive standard, and a communications playbook running a dedicated program channel where observers sign up to watch sessions live.
What keeps a small-n round honest
Seven to fifteen sessions can say a lot and can also overstate what they saw. The program's mechanics exist to keep the two apart.
- Returner re-tests. Round 3 deliberately re-recruited two Round 2 participants to react to a concept that had been revised because of what Round 2 found. Both confirmed, unprompted, that the revisions worked. That is design iteration validated by the same customers who flagged the problem.
- A change ledger between rounds. Each readout documents what product changed in response to the prior round, so readouts report movement rather than snapshots.
- Corroboration against the corpus. New findings are checked against the existing research repository before they are filed, so a round's conclusion rests on more than its own n. One launch requirement in Round 3 was settled by citing a prior survey of 527 customers rather than the round's nine sessions.
- Same-day analysis and an accuracy audit. Every session is analyzed the day it runs, findings carry traceable IDs, and a pre-ingest accuracy audit with a corrections register closes out each round before anything enters the repository.
- Named owners and observers in the room. Each concept in a round is briefed in by a named product or design owner, and owners are invited to watch their own sessions live. Product teams compete for slots.
What it has fed
- A cross-product opt-in flow was revised after Round 2, and the revision was validated by Round 3 returners: the full loop from finding to change to confirmation, inside six weeks.
- An in-flight concept was postponed before fielding, pending work the round showed it depended on. Research timing input that stopped a team from testing something that was not ready.
- A statistics concept was redirected to a different surface than the one proposed, after six of nine participants independently relocated it there. That changed where the feature should live, not just how it should look.
- A live-content concept's launch requirement (curation, not algorithm-only) was established by corroborating the round against the larger prior survey.
- A recurring support gap for high-value customers, raised unprompted three rounds running, was escalated to the service organization with a named owner. The program surfaces signal outside its own scope.
- Product leadership now feeds questions directly into rounds, from cross-round evidence pulls on competitive questions to recruiting criteria for the next round. It has become a standing channel for leadership questions, not only product-team ones.
Every round's findings also flow into the governed repository, where the self-serve Slack agent and the intake agent draw on them. The program and the AI practice compound each other: one keeps the evidence fresh, the other keeps it queryable.
What I'd tell another researcher
A program is a promise, not a study. The research design is the easy part; the design problem that matters is making the cadence indestructible, which means one calendar commitment, a runbook nobody has to reinvent, an intake rule that absorbs small asks instead of fighting them, and rigor mechanics that let a small round say something true. If your product ships continuously, your evidence should too.
Experimental usability · Dec 2025
Designing to Prevent Costly Errors: An Experimental Usability Study of Betslip Promotions
A two-test experimental study across three prototypes that identified the only layout to perform on both tasks while structurally preventing accidental spends of rewards currency
- Role
- Sole researcher: designed both tests, built the counterbalancing plan, analyzed and reported
- Timeline
- Dec 2025
- Methods
- Within-subjects interaction test (Latin Square, n = 32) · between-subjects think-aloud (n = 30) · task success and error-rate analysis · desirability scale

The problem
In the betslip, users can apply a free profit boost or one purchased with rewards currency. If a design leads someone to spend rewards currency when they meant to use a free boost, that is an accidental purchase the user did not intend, and a direct harm to trust. The team had three candidate layouts and needed to know which best helped users find and apply the correct boost while preventing that error.
The design: rigor that shipped in two weeks
Two complementary tests, run simultaneously on an unmoderated platform:
- Interaction test (quantitative), n=32, within-subjects with Latin Square counterbalancing. Every participant saw all three prototypes; the Latin Square controlled for order effects given the platform's randomization limits. Captured task success, first-click accuracy, errors, and comprehension (did the user know which boost type they had selected?).
- Think-aloud test (qualitative), n=30, between-subjects, 10 per prototype. Clean first impressions of a single prototype each, with comparison screenshots shown only at the end to gather cross-prototype preference without contaminating initial reactions.
- Non-directive task wording by design. Task 1 said only "apply a 25% profit boost," never "free." That tests whether the design itself communicates the free-vs-paid distinction, so accidental rewards-currency selections surface as the design flaw they are.
- Also piloted a standardized desirability scale for reuse across future work.
What I found
Scroll position dominated behavior. Whichever boost appeared first was selected at near-perfect rates. Prototype C (free first) hit 97.6% on the free task but only 66.7% on the paid task; Prototype B (paid first) hit 86.4% on the paid task but 57.1% on free. Each single-scroll layout optimized one task at the other's expense.
Only the separated layout (Prototype A) performed well on both (66.7% free / 77.5% paid), the balanced choice.
Error severity was the real story. In Prototype B, 4.8% of users were charged rewards currency without realizing it: an accidental purchase. In Prototypes A and C, that harmful-error rate was 0%. Prototype A's separation gave users clarity even when they erred: those who selected the paid boost knew it was the paid option.
Preference matched safety. 68.8% ranked Prototype A first in the interaction test and 53.3% chose it in the think-aloud test, citing the clear separation. "I don't want to accidentally spend my rewards" captured the sentiment.
A separate applied-state problem surfaced. The checkmark was ambiguous (confirm vs. already applied), causing users to re-tap. I recommended a distinct applied indicator and a clear removal control.
The recommendation and why
Ship Prototype A's separated layout. It is the only design balanced across both tasks, it structurally prevents accidental rewards-currency purchases, and it is the design users prefer. When the safest option is also the most preferred, the trade-off conversation gets easy. The redesign scored +1.16 composite on the piloted desirability scale (−3 to +3).
What I'd tell another researcher
This is where methodological care pays off directly. A single-prototype test, or directive task wording, would have hidden the accidental-purchase risk entirely; it only appears when you let people act on an ambiguous design and then check whether they understood what they did. Designing the study to expose the harmful error, not just measure task success, is what made the recommendation trustworthy. It also ties to responsible gaming: preventing unintended spending is a user-protection outcome, not just a usability one.
Strategic reframing · 2025
Knowing When to Pivot vs. Optimize: Reframing a Live-Betting Product Bet
Strategic reframing, mixed methods, and the judgment to deliver an unwelcome answer with a better one attached
- Role
- Sole researcher: reframed the brief, designed and ran all four methods, and delivered the readout to product leadership
- Timeline
- Jul – Aug 2025
- Methods
- Contextual inquiry during live games · competitor walkthroughs · environmental mapping · behavioral triangulation against betting history

The situation
Leadership at a top-5 US sportsbook had committed engineering resources to a pre-packaged live same-game parlay product. The ask that reached research was a validation study: help make this feature succeed. Engagement with complex live products was low, and the feature was the bet to fix it.
I did not think the question was right. Users on the platform were already betting live and already building parlays. The interesting problem was why people who did both things separately consistently avoided combining them.
The reframe
Instead of "How do we make this feature successful?", the study asked three things:
- Why do users who actively engage in both behaviors avoid combining them?
- What environmental and systemic barriers shape that avoidance?
- How should the organization tell a pivot from an optimization problem?
That reframe changed the study from tactical feature validation into a question about product direction, which meant it had to be designed to survive a room that had already made a commitment.
Method
A conventional usability test would have validated the feature in isolation and missed the context that decided the outcome. I combined four approaches in a two-week window:
- Contextual inquiry during live games, observing real betting decisions as they happened.
- Competitor walkthroughs, having participants demonstrate the same workflow on rival platforms.
- Environmental mapping, documenting where and on what devices live betting actually occurred.
- Behavioral triangulation, comparing what participants said with their own betting history.
Before fielding, I pre-aligned decision criteria with product leadership: an adoption signal below an agreed threshold would trigger a strategy conversation rather than an optimization one. Agreeing on the bar before the data arrived is what let the findings land as evidence instead of opinion.
What we found
Learned helplessness, not a usability gap. Market suspensions during live play did more than delay bets. They taught users to stop trying. The theme "I don't even bother at this point" surfaced unprompted in 8 of 12 sessions, and participants could not identify any pattern in when markets would lock. When a system is unpredictable, users stop attempting the behavior it depends on.
Context caps complexity. Only about half of live betting happened in a setting that could support a multi-step decision. Betting on the go was effectively limited to one or two legs; bar and social viewing prioritized the game over optimization. No interface change would have moved that constraint.
The hidden success story. The same users who had abandoned live same-game parlays were still building cross-game parlays live at healthy rates. "Cross-game is easy, same-game is impossible" was a consistent theme. That pointed to an immediate, low-cost optimization while the underlying infrastructure problem was addressed.
What happened
- The pre-packaged live product moved from a committed quarterly priority to an experimental track, and the engineering investment was redirected toward market stability and cross-game parlay improvements.
- The size of the realistic opportunity was reset to reflect where and how people actually bet live, which reshaped how the team sized future bets on complex live products.
- The multi-method, context-first approach became the organization's model for validating high-investment features, and product managers were asked to read the findings before their next planning cycle.
What I'd tell another researcher
The most valuable research does not always validate a direction. The judgment call was not "is this feature usable" but "is this the right question," and the craft was in designing a study that could answer the bigger question credibly inside a two-week window. Pre-aligning decision criteria, building the coalition with design partners before product and engineering, and arriving with an alternative path made it possible to deliver an unwelcome answer without burning the relationships that make research useful the next time.
Field research · 2024
Field Research Under Real Stakes: Mapping Friction Across the Betting Journey
In-person usability testing and ethnographic observation with real money on the line, synthesized into a journey-level prioritization framework that shaped a quarterly roadmap
- Role
- Designed the study and its field protocol; moderated alongside colleagues at a combined in-person event
- Timeline
- Sept 2024
- Methods
- Moderated in-person sessions during live events · ethnographic observation · journey mapping · impact × frequency prioritization

The problem
The product organization could see engagement dropping across the conversion funnel but could not agree on which friction points mattered most. Dozens of small issues were being fixed as symptoms, and the roadmap needed a defensible answer to "which of these actually costs us, and in what order?"
Method
Remote studies flatter interfaces. To see behavior that matters, I designed the study for the field: moderated in-person sessions during live events, where participants placed real bets with real consequences, combined with ethnographic observation of how they used the app alongside everything else competing for their attention. Sessions were structured around the full journey rather than a single feature, so friction could be traced upstream to its cause instead of catalogued where it surfaced.
Synthesis produced a ten-category framework scoring each issue on business impact and user frequency, which is what let three product teams with different backlogs agree on one order.
What we found
- Time-sensitive decisions broke first. 83% of participants ran into data-visualization problems at exactly the moments that required a fast decision, and the live experience rated lowest of any journey stage (5.7 of 7). Several participants said unprompted that a system lock at a key moment would send them to a competitor's app.
- Findability compounded everything downstream. 67% expressed frustration finding content, and observation showed people abandoning multi-step flows rather than pushing through navigation friction.
- Users were building their own workarounds. Progress-tracking complaints appeared in 75% of sessions, and participants were using screenshots and outside apps to track what the product should have shown them, a reliable signal that a system, not a screen, is failing.
- Fixing the cause beat fixing the symptoms. Tracing issues upstream showed that three root problems accounted for twelve of the observed downstream complaints.
What happened
- Performance and stability work at the live-decision moments moved to the top of the next quarter's roadmap across three product teams, with dedicated engineering allocation.
- The journey-level framing carried into the sportsbook research I now design and lead within the company's recurring multi-day in-person research program: sessions structured around the whole betting journey rather than a single feature.
What I'd tell another researcher
The field is where the interesting failures live. A participant in a quiet remote session will politely complete a task that the same person at a crowded bar during a fourth-quarter drive will abandon in seconds. If the decisions your product supports happen under pressure, test under pressure.
About
Hi, I'm James.
Staff UX Researcher with a Ph.D. in social psychology. I run the full method range on product timelines, and I build the practice around the research so insight compounds instead of sitting in decks.
The short version
I'm the sole UX researcher for the sportsbook product line at Fanatics Betting & Gaming, where I set the research agenda across the app and partner with the VP of Product and Director of Design on what is and isn't worth studying. Before that I built the research practice from zero at PointsBet, and before that I led institutional research functions at two colleges, where improving student retention turned out to be a UX problem with higher stakes than most apps.
The arc
- Started
- Studying decision-making and human behavior at Rutgers, running a lab of 30–40 research assistants and publishing in The Journal of Social Psychology.
- Evolved
- Directing institutional research at Glenville State and Albright, as the primary technical contributor on institutional data: SQL against the warehouse, dashboards, and mixed-methods studies that shaped retention strategy.
- Pivoted
- Bringing that toolkit to digital products at PointsBet, then Fanatics: field research, experimental usability, survey programs, and measurement infrastructure.
- Now
- Building the research practice around the research: a governed 2,100+ document repository, a self-serve Slack research agent, an AI intake agent for concept prioritization, a standing always-on research program, and the standards that make AI-assisted answers trustworthy.
How I work
Evidence should lead, not validate. I start from the decision that needs to be made and work backward, run parallel streams instead of sequential phases, and aim for "good enough to decide" rather than academically perfect. The PhD matters less for the credential than for knowing when each method is appropriate, what threatens its validity, and how to design around that on a two-week timeline.
The AI part is a governance problem as much as a tooling one. Every AI-assisted output in my practice runs through accuracy verification against source, provenance on every claim, and recency down-weighting before it reaches a stakeholder. Fast is only useful if it stays honest.
Beyond the work
Former touring musician: bass for Hand Me Down Buick, with Warped Tour and House of Blues on the résumé, and stakeholder management learned the hard way in 2 AM venue negotiations. Father of two boys who are my toughest critics. Pennsylvania resident. Toastmasters alum.
What I care about
Teams that value strategic thinking over tactical execution, evidence over opinion, and outcome measurement over output tracking. If any of the work here is useful to you, or you want to compare notes on research practice, say hello.
How I work
Programs, instruments, and systems, not just studies
Individual studies answer individual questions. This page is about the things I build so research compounds: standing programs, measurement infrastructure, and the AI-augmented practice that keeps it all queryable.
In-person research with real stakes
Fanatics runs a recurring multi-day, in-person research program: 25+ moderated sessions per event with real customers, multiple moderators working in parallel. The program was here before me; what's mine is the sportsbook research inside it. I design the discussion guides and testing topics, lead the sportsbook sessions, and turn the output into the shared source layer that downstream workstreams draw on, from promotions to social and personalization. It is a collaborative effort across the insights organization, and that is part of what makes it work.
The field research is my own design: observing bettors in their natural context during live games. Remote studies flatter interfaces; when you watch someone bet at a crowded bar during a fourth-quarter drive, you see what actually happens. Case study →
Always-on research: a standing evidence pipeline
Most research is commissioned one study at a time, and it thins out exactly where products need it most, after launch. I built a standing program instead: moderated rounds with our highest-value customer segment every three weeks, alternating between post-launch evaluation, discovery, and concept review as product needs dictate. A fixed runbook turns a product team's Monday question into a customer-evidence readout inside two weeks: intake, warehouse-based recruiting, sessions analyzed the same day they run, synthesis and shareout by Friday.
The design compounds. Participants return in later rounds to re-test designs that changed because of what they said, in one case confirming, unprompted, that the revisions worked. A change ledger tracks what product shipped between rounds, so readouts report movement, not snapshots. Product managers and designers brief their concepts into rounds and watch their own sessions live. And every round's findings flow into the governed repository, where the self-serve agents can draw on them. Four rounds and 40+ sessions in the first three months, feeding decisions from concept iteration to launch requirements to one postponement that stopped a team testing something that wasn't ready. Case study →
Co-creation: customers in the room
I designed customer roundtables that bring actual users to product offsites. Not behind a one-way mirror: in the room, with designers, PMs, and engineers asking questions directly. When stakeholders hear it themselves, you skip the "convince leadership" step entirely.
Controlled experimentation: rigor that ships
Three prototypes, Latin Square counterbalancing to control for order effects, 62 participants across two parallel tests, task success compared quantitatively. Clear winner, defensible methodology, decision made in two weeks. Case study →
Measurement infrastructure: systems, not just studies
I created the measurement instruments my team now uses across studies. Individual studies are valuable; repeatable measurement systems that compound over time are transformational.
An integrated priority-ranking framework combines these scores with qualitative signal and strategic alignment into composite scores that give stakeholders a clear, defensible order. The same feature-prioritization method (rating plus ranking, with a data-quality gate that removes straight-line responders before anything is ranked) has now run three times: a six-feature personalization cycle, a thirteen-feature promotions cycle at twice the scale (636 raw responses, 503 after a 20.9% quality cut), and a parlay-features cycle.
The AI-augmented practice
A governed research repository, five AI research roles (a self-serve Slack agent, an intake agent for leadership prioritization, an opportunity-backlog role that tracks off-plan needs across studies, a research-and-strategy prioritization role, and repository Q&A and briefs), and seven automated workflows keeping the corpus current, all held to researcher-grade standards: attributed verbatims, evidence dated per finding, explicit gaps, an accuracy gate. Case study →
The full methodological range
Quantitative: survey design with psychometric validation, experimental and quasi-experimental design, data-quality auditing, statistical analysis (SEM, factor analysis, regression, ANOVA) in Python and R. Qualitative: in-depth interviews, contextual inquiry, ethnographic observation, moderated and unmoderated usability testing, diary studies, and standing recurring research programs. Mixed: triangulation across qual and quant, qual→quant sequencing, integrated frameworks that combine multiple data streams.
Practice building: three times
Glenville State and Albright College: director-level roles creating institutional research functions where none existed. PointsBet: sole researcher to a full practice with repository, taxonomy, and stakeholder rituals. Fanatics: an inherited practice elevated with AI tooling, measurement infrastructure, and standard frameworks, positioning research as a strategic partner rather than a validation service.
Contact
Say hello.
Questions about the work, the methods, or the AI-augmented practice, or just want to compare notes on research: I'd like to hear from you.
I typically respond within a day or two.