Cohort Identification & Feasibility
Translate a plain-English protocol into a defined patient cohort with accurate counts, so you know what you have before you commit.
You're ready to go, start your request directly.
Submit RequestThe Clinical Data Core (CDC) is a service team within the McWilliams School of Biomedical Informatics at UTHealth Houston. We support investigators across the institution who need clinical data for research — cohort identification, feasibility counts, and de-identified extracts — and we treat every request as a research partnership, not a ticket in a queue.
Our work centers on one idea: get researchers a dataset they can trust and defend, faster. We do that by pairing experienced informatics staff with a structured, auditable extraction process so the data that reaches you is accurate, reproducible, and fully within IRB and HIPAA guardrails.
Our retrieval workflow is built to reach a first reviewable artifact quickly, then refine it with your feedback. A few recent proof points:
Speed claims are deliberately scoped to the first reviewable artifact, not final study completion. Comparators are user-provided.
We offer row-level PHI data or datasets of counts only, from Memorial Hermann source data from Clarity and Cerner.
Translate a plain-English protocol into a defined patient cohort with accurate counts, so you know what you have before you commit.
Governed, reviewable datasets scoped to the minimum necessary for your study, with provenance and evidence labels attached.
Same-day feedback cycles with the investigator. We refine cohort logic against real EHR patterns through embedded review, not handoffs.
We maintain a HIPAA accounting of disclosures — a record of who received which data — and track custody of every field.
Pre-screening, recruitment feasibility, and eligibility cohorts for trial teams — identify candidate patients against protocol criteria before activation.
Stand up and maintain disease, device, or outcomes registries — reproducible cohort definitions, longitudinal updates, and a documented audit trail.
CareFlow is our agentic AI retrieval engine. Think of it as a bonded courier service: the request is carried through audited checkpoints, and what comes back is a sealed, traceable parcel — never the keys to the warehouse. The AI agent plans and drafts the query, but it never touches patient identifiers, and a human always reviews the result before it is released.
The engine follows a one-directional waterfall method — metadata samples frozen cohort extract audit deliver . Because the cohort is frozen before any data is pulled, the same protocol always yields the same cohort, and the audit trail can prove it.
Plain-English protocol parsed into atomic criteria
Each atom queried alone against a read-only probe — no PHI
Cached results combined by set logic into one cohort
Counts, provenance & logic confirmed before anything is released
Governed, reproducible extract with a full audit trail
You stay in the loop at every step — confirm criteria, review counts, and request revisions. CareFlow never advances a stage without a human go-ahead.
A mirrored process. CareFlow doesn't replace our analysts — it runs alongside them. Every automated step has a human counterpart, so the investigator is engaged continuously rather than waiting on a black-box handoff:
How a research request moves from submission to delivered dataset — five steps, the way a PI sees them.
Where it starts
PI obtains IRB & MH CIRI approval and logs a request in ServiceNow. CDC team receives notification and acknowledges receipt.
PI SubmitsWe engage
CDC team engages with PI to approve/confirm scope of work and drafts an MOU. PI reviews MOU, signs/submits to CDC Team. Request moves through the CDC Queue.
PI EngagedWork in progress
CDC developer builds the dataset extract from the requested sources. PI receives a status update that work has started.
Status: CodingYou validate
PI reviews the dataset against acceptance criteria. Any rework triggers a quick loop back — never a restart. Iterations are captured.
PI Signs offYou have your data
Final dataset link delivered through SNOW with REQ-ID, PI, date/time, and link to data. Closure email includes a satisfaction survey.
Dataset LiveGovernance is part of the retrieval layer itself, not bolted on afterward. There are five guards between any external team and Memorial Hermann data — no external team touches the retrieval engine directly.
Credentialed requesters only; IRB & CIRI, Data Use Agreements verified upstream/downstream.
Every request is scoped to the question and time-bound — access expires at study end.
A CDC-operated retrieval engine sits between the requester and the protected core.
Results are screened for re-identification risk before anything is released.
Every action is logged, immutable, and reviewable on demand — running beneath every guard.
Studies using Memorial Hermann clinical data require a Memorial Hermann Clinical Innovation & Research Institute (CIRI) submission. You can begin that process here:
Tell us the clinical question behind your study — not just a row count — and we will help you scope it.
How our services are priced
The Clinical Data Core operates as a cost-recovery service center under federal Uniform Guidance (2CFR 200), like other UTHealth core facilities. Our services are billed at a single published internal rate of $150 per hour, which is direct-chargeable to sponsored awards and grant budgets. External rates are available on request. Your estimate is always free Every internal UTHealth request begins with two complimentary hours of consultation and scoping. Before any billable work starts, you’ll receive a written quote with estimated hours and cost – and no work is charged until you’ve approved the scope and confirmed a funding source. If scope changes, we pause and re-quote (if needed) before continuing. Contact the Clinical Data Core Meet the Team