Evidence Management

How to Build an Evidence Matrix From PDFs, Web Pages, and Interviews

Build a traceable evidence matrix across PDFs, web pages, interviews, videos, and notes. Use practical columns and a verification workflow for defensible analysis and cited deliverables.

People in the office

Updated October 9, 2026.

What an evidence matrix does in mixed-source research

To build an evidence matrix, define the research question, decide what one row represents, extract evidence consistently, preserve the source location, and review contradictions before drawing conclusions. The point is not to make a prettier note dump. It is to connect each claim in the final analysis back to the PDF page, web passage, interview line, video timestamp, or analyst note that supports it.

An evidence matrix is a structured record of claims, findings, source passages, provenance, and quality judgments. It is useful when a deliverable has to survive review: a policy brief, litigation memo, market scan, medical evidence summary, due diligence report, grant narrative, or research synthesis.

It is not the same as a bibliography. A bibliography tells you what was consulted. A matrix tells you what each source contributed, where the evidence lives, how strong it is, and what still does not fit.

Evidence matrix linking claims to mixed research sources

A simple policy example makes the difference clear. Suppose the question is: “What evidence supports expanding remote monitoring for chronic disease management in rural clinics?”

The source collection might include:

  • A PDF evaluation report from a health agency

  • A government web page with current reimbursement rules

  • An interview transcript from a clinic director

  • A recorded stakeholder meeting

  • Analyst notes from a review of implementation barriers

A weak workflow stores these in five places and drafts from memory. A better workflow gives each item a stable source ID, extracts the relevant passage, records the page or timestamp, labels whether the evidence is empirical, regulatory, experiential, or interpretive, and flags disagreements before the brief is written.

A mixed-source evidence matrix can support professional analysis, but it does not replace a formal systematic-review protocol. If the project requires formal screening, risk-of-bias assessment, registered methods, or meta-analysis, the matrix should sit inside that protocol rather than pretend to be one.

Start with the research question and decide what each row represents

The matrix will sprawl unless the research question is narrow enough to exclude interesting but irrelevant material. A useful question names the decision, population, setting, concept, or outcome being examined. The stepwise research-question literature makes the same basic point: start with a subject of interest, do preliminary work, then define the specific question before collecting and interpreting evidence (Ratan et al., 2019).

For mixed-source work, a research question often needs subquestions. In the rural clinic example, the top-level question could break into:

  • What outcomes have been reported for remote monitoring?

  • What reimbursement or regulatory constraints apply?

  • What implementation barriers do clinic staff report?

  • What evidence is specific to rural settings?

  • Where do stakeholder claims conflict with published findings?

Then decide what each row in the matrix represents. This is the most important design choice.

Common row units include:

  • Claim-level row: one discrete claim supported by one passage or source location.

  • Finding-level row: one finding that may be supported by several sources.

  • Source-level row: one row per source, with columns for themes.

  • Participant-level row: one row per interviewee, stakeholder, organization, or case.

  • Case-level row: one row per site, jurisdiction, company, policy, intervention, or event.

  • Theme-level row: one row per analytic theme, with evidence grouped beneath it.

  • Comparison-point row: one row per criterion, such as cost, access, risk, feasibility, or outcome.

For traceable deliverables, claim- or finding-level rows usually work best. They force each assertion to carry its own provenance. Source-level rows are faster for early reading, but they can hide the exact passage behind broad summaries.

If the project is limited to academic literature, a literature matrix may be enough. For that narrower workflow, see Otio’s guide to comparing sources in a literature review. A mixed-source evidence matrix needs extra care because PDFs, web pages, interviews, recordings, and notes do not have the same evidentiary status.

Define the inclusion rule before extraction begins. For example:

  • “Add one row for each discrete finding that answers a subquestion.”

  • “Add one row for each source’s position on each implementation barrier.”

  • “Add one row for each stakeholder claim that will be cited or challenged.”

  • “Add one row only when the evidence has an exact passage and location.”

Separate descriptive rows from interpretive rows. A descriptive row says, “The clinic director reported that device setup took 30 minutes per patient.” An interpretive row says, “Setup burden may reduce adoption in understaffed clinics.” Both may belong in the matrix, but they should not look identical.

Use columns that preserve traceability and context

An evidence matrix can live in Excel, Google Sheets, Airtable, Notion, a database, or a structured document. Brandeis University’s writing resources note that matrix-method literature reviews can be built in tools such as Excel, Word, OneNote, Google Sheets, or Numbers (Brandeis University). The tool matters less than the discipline of the fields.

Use enough columns to let another reviewer inspect the chain from conclusion back to source.

Core columns:

  • Source ID: a stable label such as PDF-03, WEB-07, INT-02, VID-01, NOTE-04.

  • Source title or label: report title, page title, participant label, meeting name, or note title.

  • Source type: PDF report, peer-reviewed article, government web page, interview transcript, video, audio, observation note, analyst memo.

  • Author, publisher, speaker, or participant ID: whoever produced the evidence.

  • Publication, interview, recording, or collection date: the date attached to the source.

  • Version or access date: especially important for web pages, updated PDFs, transcripts, and living documents.

  • Research question or theme: the subquestion the row addresses.

  • Claim or finding: the concise statement the evidence supports.

  • Supporting quotation or passage: the exact text, faithful transcript excerpt, or visual description.

  • Location or timestamp: page number, section heading, paragraph, transcript line, timestamp, slide number, figure, table, or URL anchor.

  • Strength or relevance: why the row matters and how directly it answers the question.

  • Contradictions or caveats: conflicting sources, limitations, exceptions, or scope issues.

  • Follow-up action: verify source, request transcript approval, check later version, find corroboration, downgrade claim, or exclude.

Optional fields help when the project is larger or regulated:

  • Participant or case ID

  • Method used by the source

  • Population or sample

  • Geographic scope

  • Jurisdiction

  • Sector or organization type

  • Reviewer

  • Confidence rating

  • Source status: included, excluded, provisional, superseded

  • Link to original file, page, recording, or transcript

  • Anonymisation or confidentiality requirement

Evidence matrix columns for traceability and quality control

The supporting excerpt is the column people are most tempted to skip. Do not skip it. A summary alone forces every later reviewer to trust the extractor’s interpretation. A short passage lets the reviewer see whether the claim is actually supported, overstated, or missing a qualification.

A completed row might read like this in prose:

Source ID WEB-03 is a government reimbursement page accessed on October 9, 2026. It addresses the “payment constraints” subquestion. The extracted passage says remote monitoring reimbursement requires specific documentation and eligible service conditions. The matrix row records the page title, URL, relevant heading, access date, and exact passage. The supported finding is: “Reimbursement may be available only when documentation and eligibility requirements are met.” The caveat is that the page may change, so the row is marked for verification before final publication.

Keep source identifiers stable. If a web page changes, WEB-03 should remain WEB-03, with the version or access date updated. If a PDF is replaced by a later report, keep both versions separate unless the older one is formally excluded.

Standardize evidence across PDFs, web pages, interviews, video, and notes

Standardization does not mean pretending every source type is equivalent. It means each source can be inspected with the same basic questions: who said it, when, where, in what context, and with what limits?

For PDFs, record:

  • Document title

  • Author or issuing organization

  • Publication date

  • Version date, revision date, or file date when available

  • Page number

  • Section heading

  • Figure, table, appendix, or footnote number when relevant

  • Whether the text came from selectable text or OCR

  • Link or file path to the stored document

OCR matters because scanned text can introduce errors. If the evidence comes from a scanned PDF, verify the passage manually against the visible page before quoting it.

For web pages, record:

  • Page title

  • Publisher or organization

  • Author if listed

  • Publication date, update date, or “no date”

  • Access date

  • URL

  • Relevant heading or page section

  • Extracted passage with enough surrounding context

  • Any sign that the page is dynamic, periodically updated, or jurisdiction-specific

A current URL is not a stable claim. It is a path to a page that may change. The access date and excerpt matter.

For interviews, record:

  • Participant or interview ID

  • Interview date

  • Interviewer if relevant

  • Transcript version

  • Transcript line number, paragraph, or timestamp

  • Whether the quote is approved, anonymised, or restricted

  • Whether the entry is a direct quote or researcher paraphrase

Treat an interview as evidence of what the participant reported, experienced, believed, or observed. Do not automatically treat it as proof that the reported event occurred exactly as described or generalizes beyond that participant’s position.

For audio and video, record:

  • File or platform source

  • Recording date

  • Speaker if identifiable

  • Timestamp start and end

  • Whether the evidence is spoken dialogue, on-screen text, visible behavior, slide content, or analyst interpretation

  • Transcript status if one exists

  • Link to the stored file or page

For projects with lectures, meetings, or field recordings, a summarizer can help with triage, but summaries are not evidence-ready by default. Otio’s post on Google Drive video summarizer tools is useful if recordings are part of the source set and need to become searchable notes before verification.

For researcher notes, label the note type clearly:

  • Observation

  • Hypothesis

  • Decision

  • Interpretation

  • Question

  • Method note

  • Meeting note

Do not treat analyst notes as independent source evidence unless the note records an observation from the research process. A note that says “rural clinics likely face staffing barriers” is an interpretation. A note that says “during the site visit, two staff members reported that device setup was handled by one nurse” is closer to observational evidence, but still needs context and limitations.

The failure mode is subtle: once everything appears in the same grid, it starts to feel equally weighted. A peer-reviewed study, a stakeholder statement, a web claim, and an analyst hypothesis can sit beside each other, but the matrix should make their differences visible.

Extract evidence while preserving the original passage and location

Extraction should be boring and repeatable. That is the point.

Use this sequence for each source:

  1. Read, watch, or listen for material relevant to a subquestion.

  2. Capture the exact passage, transcript excerpt, or faithful description.

  3. Record the location immediately.

  4. Add a neutral description of what the source says.

  5. Write the claim or finding the passage supports.

  6. Add caveats, contradictions, or follow-up checks.

  7. Confirm that the original source is still accessible.

The University of North Carolina’s evidence-synthesis guidance describes extraction and quality assessment as the next stage after selecting included sources, with data extracted from each source and assessed for quality (UNC Libraries). Even when the project is not a formal evidence synthesis, the same discipline applies: extract first, judge after.

Capture enough surrounding context. If a report says, “Remote monitoring reduced emergency visits among high-risk patients in the pilot, but the evaluation did not include clinics without broadband support,” the matrix should not reduce that to “remote monitoring reduced emergency visits.” The condition and limitation are part of the evidence.

For PDFs, record the page number as soon as the passage is captured. Add section headings because page numbers can shift across versions. If the evidence is in a figure or table, record the figure or table number, not just the page.

For web pages, copy the relevant heading and access date with the passage. If the page is a policy page, product page, or government guidance page, assume it may change.

For interviews, preserve the participant’s wording inside quotation marks. Do not silently clean up grammar, remove hedging, or smooth a quote into the language you wish they had used. If a quote needs light editing for readability, mark that decision in the matrix or use paraphrase instead.

For video and audio, use timestamp ranges rather than a single timestamp when the relevant evidence spans a statement. A range such as 00:14:22–00:15:10 is more useful than “around 14 minutes.”

Store the original source beside the matrix entry. A spreadsheet full of citations is brittle if the files live in personal downloads folders, private email threads, or expired links. The matrix is the index; the source library is the evidence base.

[[OTIO_INLINE_PROMO:%7B%22title%22%3A%22Can%20you%20trace%20every%20excerpt%20to%20its%20source%3F%22%2C%22description%22%3A%22Add%20your%20PDFs%2C%20web%20pages%2C%20interviews%2C%20and%20notes%20to%20one%20Otio%20library%2C%20then%20use%20cited%20answers%20to%20locate%20candidate%20passages%20before%20checking%20each%20page%20or%20timestamp.%22%7D]]

Compare evidence without flattening differences in quality or method

Do not score evidence before recording what it says. First extract the passage. Then assess relevance, strength, and limitations.

Useful bases for a strength or relevance judgment include:

  • Directness: Does the source answer the exact question, or only a related one?

  • Methodological fit: Was the evidence collected in a way that supports the claim?

  • Recency: Is the source current enough for the decision?

  • Scope: Does the population, jurisdiction, organization, or setting match the project?

  • Corroboration: Is the point supported by independent sources?

  • Potential bias or incentive: Does the source have a stake in the claim?

  • Transparency: Are methods, definitions, data, or assumptions visible?

  • Precision: Does the source provide specific evidence or only a broad assertion?

Keep separate fields for direct evidence, interpretation, and confidence. A persuasive quote from a senior stakeholder may be important, but it is not the same as a measured outcome. A recent web page may be operationally relevant, but it may not have the evidentiary weight of a well-designed evaluation.

Record contradictions explicitly. Sources can disagree in different ways:

  • Fact disagreement: one says a policy exists; another says it was repealed.

  • Definition disagreement: sources use the same term differently.

  • Estimate disagreement: costs, prevalence, risk, or effect sizes differ.

  • Causal disagreement: sources disagree about why an outcome occurred.

  • Perspective disagreement: stakeholders value or interpret the same facts differently.

  • Scope disagreement: one claim applies nationally; another applies only to a region or subgroup.

Do not average contradictions away. A single confidence score can hide the thing that matters most: the sources were collected through different methods, under different incentives, for different purposes.

A government web page, for example, may be the best source for current eligibility rules. It may be a poor source for whether clinics can implement those rules without added staff. An interview may reveal implementation burden, but it may not establish how common that burden is across sites.

The reviewed matrix should show both.

Run a quality-control pass before drafting conclusions

A matrix is not ready for drafting just because every row has text in it. Run a quality-control pass before conclusions harden.

Check that every important row has:

  • Source ID

  • Source type

  • Date, version date, or access date

  • Exact passage, faithful transcript excerpt, or clear visual description

  • Page number, heading, line number, timestamp, or other location

  • Link or path back to the original source

  • Research question or theme

  • Claim or finding

  • Caveats, contradictions, or quality judgment when relevant

Then look for failure patterns:

  • Interpretation with no direct evidence

  • A quote that no longer supports the claim beside it

  • Claims that rely on a single interested source

  • Web claims without access dates

  • PDF excerpts without page numbers

  • Interview paraphrases that are presented as quotations

  • Duplicated evidence counted as independent corroboration

  • Broad conclusions drawn from narrow settings

  • Missing contradictory rows that were inconvenient but material

PRISMA 2020 notes that presenting key study characteristics in a table or figure can help comparison across studies (Page et al., 2021). The same principle applies outside formal reviews: comparison works only when the key characteristics are visible enough to challenge.

Ask a second reviewer to inspect the matrix if the stakes justify it. If no second reviewer is available, use a different AI model as a challenger, not as an authority. Ask it to identify overclaims, missing caveats, mismatched citations, alternative interpretations, and rows that should be downgraded. Then check every challenge against the original sources.

Before drafting, group rows by research question or theme. Label each row’s role:

  • Central evidence

  • Contextual evidence

  • Contradictory evidence

  • Weak but suggestive evidence

  • Stakeholder perspective

  • Analyst interpretation

  • Provisional or unresolved

Unresolved contradictions are not embarrassing. They are often findings. A brief that says “evidence is mixed because published evaluations report improved outcomes while clinic interviews identify staffing barriers that may limit adoption” is stronger than one that forces agreement where none exists.

How Otio can support the evidence-matrix workflow

Otio’s AI PDF reader can support the source side of this workflow: PDFs, web pages, interview audio or transcripts, videos, notes, and other files can sit in one research library instead of being scattered across tabs, downloads, cloud folders, and chat sessions.

That does not mean the matrix has to live inside Otio. Many teams should keep the working matrix in a spreadsheet, database, or structured document because sorting, filtering, reviewer assignment, and version control may already be built around that format. Otio’s role is the permanent multi-source workspace beside the matrix.

A practical workflow looks like this:

  1. Add the PDF reports, web pages, transcripts, recordings, and notes to one project library.

  2. Ask a cited question across the collection, such as: “Which sources discuss staffing burden for remote monitoring?”

  3. Open the cited source passages rather than relying on the answer alone.

  4. Save useful selections to notes when they need further review.

  5. Transfer only verified excerpts, locations, and claims into the matrix.

  6. Use a second model to challenge the synthesis or look for contradictions.

  7. Return to the original source before citing anything in the deliverable.

Otio supports per-chat model selection across major model families, which is useful when a synthesis needs a second pass. A model that writes a clean summary may miss caveats; another may be better at adversarial critique. The point of using multiple AI models is not to vote on truth. It is to expose weak interpretations before a human reviewer signs off.

This is especially useful in academic research workflows, where PDFs, notes, and citation context need to stay together, and in policy research workflows, where mixed evidence often feeds briefs, consultation responses, and recommendations.

Keep the boundary clear. Otio can organize and interrogate the source collection, surface cited passages, help save selections into notes, and support cross-checking. It does not automatically produce a methodologically valid evidence matrix. It does not replace spreadsheets, qualitative-coding platforms, systematic-review software, discovery engines, expert judgment, or source checking.

Every quotation, citation, timestamp, page number, interpretation, and conclusion still needs researcher review.

Turn the reviewed matrix into a defensible deliverable

Draft from the reviewed rows, not from memory. Memory favors the most vivid source, the cleanest quote, or the last thing read. The matrix forces the draft to follow the evidence.

For each section of the deliverable, group the rows by research question:

  • What evidence converges?

  • What evidence diverges?

  • Which sources are strongest for the claim?

  • Which sources provide context but not direct support?

  • Which stakeholder perspectives should be labeled as perspectives?

  • Which contradictions remain unresolved?

  • Which claims should be softened or removed?

Citations should point to the original source and precise location. Do not cite an AI answer as a substitute for the source passage. If the final report says a PDF evaluation found a change in outcomes, cite the PDF page, table, or section. If it uses an interview quote, cite the transcript location or approved participant reference according to the project’s confidentiality rules.

Preserve the audit trail. Keep:

  • The final matrix version

  • Source versions or access dates

  • Extraction decisions

  • Excluded or downgraded rows

  • Unresolved questions

  • Review notes

  • Changes made during drafting

For a report or briefing, distinguish documented findings from stakeholder perspectives, analyst interpretation, and recommendations. Readers should be able to tell when a statement is directly supported, when it is inferred, and when it is a judgment based on the evidence.

The next practical step is simple: add the PDFs, web pages, interview audio or transcripts, videos, and notes to one Otio library, ask for a cited comparison of the evidence, then transfer only verified entries into the working matrix. That is how the workflow stays fast without giving up traceability.

FAQ

Q: Should an evidence matrix be a spreadsheet?
A: A spreadsheet is often useful for sorting, filtering, and reviewing rows, but the matrix can also live in a database or structured document. The important requirement is stable provenance: each entry should connect a claim to its source, passage, location, and quality judgment.

Q: How do you handle conflicting evidence in an evidence matrix?
A: Keep conflicting entries rather than averaging them away. Record what differs, assess each source’s method and context, and mark whether the disagreement is resolved, explainable by scope or definitions, or still unresolved.

Q: Can AI create an evidence matrix from interviews and documents?
A: AI can help locate passages, organize candidate entries, and compare themes across sources, but a researcher must verify quotations, timestamps, page numbers, interpretations, and source quality. AI output should be treated as a draft for review, not as an automatically valid matrix.

Q: What is the difference between an evidence matrix and a literature review matrix?
A: A literature review matrix usually organizes academic publications, while a mixed-source evidence matrix can include PDFs, web pages, interviews, video, audio, and notes. The mixed-source version must make source type, provenance, collection method, and evidentiary limits explicit.

[[OTIO_FOOTER_PROMO:%7B%22title%22%3A%22Bring%20your%20own%20sources%20into%20the%20matrix%22%2C%22description%22%3A%22Use%20Otio%20to%20compare%20cited%20evidence%20across%20your%20source%20collection%2C%20then%20transfer%20only%20verified%20claims%2C%20excerpts%2C%20locations%2C%20and%20caveats%20into%20your%20working%20matrix.%22%7D]]