Skip to Content
SBMI Horizontal Logo

Clinical Data Core – UTHealth Houston

UTHealth Houston Clinical Data Core

We help researchers turn a clinical protocol into a reviewable, governed dataset, in days, not months, using a modern, secure, agentic retrieval workflow called CareFlow.

Already have IRB approval & an MH CIRI approval?

You're ready to go, start your request directly.

Submit Request

Need to discuss your request?

Book time with the CDC during weekly office hours.

Book Office Hours

Who We Are

The Clinical Data Core (CDC) is a service team within the McWilliams School of Biomedical Informatics at UTHealth Houston. We support investigators across the institution who need clinical data for research — cohort identification, feasibility counts, and de-identified extracts — and we treat every request as a research partnership, not a ticket in a queue.

Our work centers on one idea: get researchers a dataset they can trust and defend, faster. We do that by pairing experienced informatics staff with a structured, auditable extraction process so the data that reaches you is accurate, reproducible, and fully within IRB and HIPAA guardrails.

Quick Turnaround, Proven

Our retrieval workflow is built to reach a first reviewable artifact quickly, then refine it with your feedback. A few recent proof points:

7 days
cCMV protocol first reviewable cohort package (Feb 17 alignment Feb 24 artifacts)
~17–30×
Faster to a first artifact than the traditional 4–7 month extraction baseline
< 3 wks
NICU cohort feasibility delivered for the Bravo study — 1,405 patients identified

Speed claims are deliberately scoped to the first reviewable artifact, not final study completion. Comparators are user-provided.

Data We Offer

We offer row-level PHI data or datasets of counts only, from Memorial Hermann source data from Clarity and Cerner.

Services We Offer

Cohort Identification & Feasibility

Translate a plain-English protocol into a defined patient cohort with accurate counts, so you know what you have before you commit.

De-identified Data Extracts

Governed, reviewable datasets scoped to the minimum necessary for your study, with provenance and evidence labels attached.

Research Partnership & Validation

Same-day feedback cycles with the investigator. We refine cohort logic against real EHR patterns through embedded review, not handoffs.

Governance & Disclosure Tracking

We maintain a HIPAA accounting of disclosures — a record of who received which data — and track custody of every field.

Clinical Trials

Pre-screening, recruitment feasibility, and eligibility cohorts for trial teams — identify candidate patients against protocol criteria before activation.

Registry Building

Stand up and maintain disease, device, or outcomes registries — reproducible cohort definitions, longitudinal updates, and a documented audit trail.

How CareFlow Works

CareFlow is our agentic AI retrieval engine. Think of it as a bonded courier service: the request is carried through audited checkpoints, and what comes back is a sealed, traceable parcel — never the keys to the warehouse. The AI agent plans and drafts the query, but it never touches patient identifiers, and a human always reviews the result before it is released.

The engine follows a one-directional waterfall method metadata samples frozen cohort extract audit deliver . Because the cohort is frozen before any data is pulled, the same protocol always yields the same cohort, and the audit trail can prove it.

  1. Convert

    Plain-English protocol parsed into atomic criteria

  2. Execute

    Each atom queried alone against a read-only probe — no PHI

  3. Assemble

    Cached results combined by set logic into one cohort

  4. Review

    Counts, provenance & logic confirmed before anything is released

  5. Deliver

    Governed, reproducible extract with a full audit trail

PI / Researcher

You stay in the loop at every step — confirm criteria, review counts, and request revisions. CareFlow never advances a stage without a human go-ahead.

A mirrored process. CareFlow doesn't replace our analysts — it runs alongside them. Every automated step has a human counterpart, so the investigator is engaged continuously rather than waiting on a black-box handoff:

CDC Engineering track (CareFlow)

  • Parses the protocol into verifiable atoms
  • Runs read-only probes & caches counts
  • Assembles the cohort by set logic
  • Logs every action to an immutable audit trail

PI / Researcher track (human in the loop)

  • Frames the clinical question & reviews atom logic
  • Validates counts against clinical expectation
  • Flags criteria for revision in same-day cycles
  • Approves the final cohort before any extract

CareFlow Workflow | Intake & Delivery

How a research request moves from submission to delivered dataset — five steps, the way a PI sees them.

  1. Request Submitted

    Where it starts

    PI obtains IRB & MH CIRI approval and logs a request in ServiceNow. CDC team receives notification and acknowledges receipt.

    PI Submits
  2. Acknowledged &
    Reviewed

    We engage

    CDC team engages with PI to approve/confirm scope of work and drafts an MOU. PI reviews MOU, signs/submits to CDC Team. Request moves through the CDC Queue.

    PI Engaged
  3. Building Your
    Dataset

    Work in progress

    CDC developer builds the dataset extract from the requested sources. PI receives a status update that work has started.

    Status: Coding
  4. Review & Iterate

    You validate

    PI reviews the dataset against acceptance criteria. Any rework triggers a quick loop back — never a restart. Iterations are captured.

    PI Signs off
  5. Delivered

    You have your data

    Final dataset link delivered through SNOW with REQ-ID, PI, date/time, and link to data. Closure email includes a satisfaction survey.

    Dataset Live

Security & Guardrails Are Built In

Governance is part of the retrieval layer itself, not bolted on afterward. There are five guards between any external team and Memorial Hermann data — no external team touches the retrieval engine directly.

  1. 1 · Identity & Approval

    Credentialed requesters only; IRB & CIRI, Data Use Agreements verified upstream/downstream.

  2. 2 · Minimum Necessary

    Every request is scoped to the question and time-bound — access expires at study end.

  3. 3 · Mediated Engine

    A CDC-operated retrieval engine sits between the requester and the protected core.

  4. 4 · Egress Review

    Results are screened for re-identification risk before anything is released.

  5. 5 · Audit Trail

    Every action is logged, immutable, and reviewable on demand — running beneath every guard.

  • Read-only access — the engine can query but never mutate the source data.
  • Aggregates and small samples only during design — no PHI in the design loop, by construction.
  • No secrets in prompts; credentials never leave the secure environment.
  • External AI tool egress is blocked by default — only reviewed, predefined skills can run.
  • Human-in-the-loop approval gates for any sensitive or irreversible action.

Working with Memorial Hermann Data?

Studies using Memorial Hermann clinical data require a Memorial Hermann Clinical Innovation & Research Institute (CIRI) submission. You can begin that process here:

Get Started

Tell us the clinical question behind your study — not just a row count — and we will help you scope it.

How our services are priced

The Clinical Data Core operates as a cost-recovery service center under federal Uniform Guidance (2CFR 200), like other UTHealth core facilities. Our services are billed at a single published internal rate of $150 per hour, which is direct-chargeable to sponsored awards and grant budgets. External rates are available on request.

Your estimate is always free

Every internal UTHealth request begins with two complimentary hours of consultation and scoping. Before any billable work starts, you’ll receive a written quote with estimated hours and cost – and no work is charged until you’ve approved the scope and confirmed a funding source. If scope changes, we pause and re-quote (if needed) before continuing.

Contact the Clinical Data Core Meet the Team