Skip to content
Data TA
FAQ

Frequently asked questions

What is Data TA?

Data TA is an experimental search app that's designed to support instructors who are looking for real-world datasets to use in class. Our goal is to allow instructors to search by topic and find datasets that have been used in published research about that topic.

Is Data TA an official ICPSR product?

Depends on what you mean by "official". Data TA is part of a research project out of Dr. Libby Hemphill's research group at the University of Michigan School of Information. Libby is also a faculty member at ICPSR.

Data TA is built on ICPSR's public metadata. ICPSR curates the data; we mapped how it has been used. Data TA never hosts data. The "Get this dataset" link takes you to ICPSR's own study page.

Where does the data come from?

Every dataset is a real ICPSR study. We built a knowledge graph from three sources: DataCite (a DOI registry), OpenAlex (an open index of scholarly papers), and the ICPSR Bibliography. We cross-check citation links across these sources rather than trust a single one. The graph holds 31,160 ICPSR records, 98,626+ publications, and 199,136+ citation links. So far, 9,290+ of those datasets (30%) have at least one citing paper.

What does "used by 47+ papers" mean? Why the "+"?

The count is a floor, not a total. Authors don't always formally cite data they use, so citation tracking misses many real uses. Informal mentions in a paper's text go untracked. Some datasets have multiple versions, and each version has a unique DOI. Citations may be spread across those DOIs. So the true number of citations is often higher than what we show. The "+" says that honestly.

The key study on each card comes from the citing papers. We prefer citations that more than one source confirms. After that, we pick the most-cited paper. We never pick a retracted paper as the key study. It is a key study that used the dataset, not the definitive paper about the dataset.

How does search work? Is it AI?

Partly. Search runs two lanes. One lane matches your words against curated topic tags and study text. The other lane matches on meaning using text embeddings. A fixed formula, reciprocal rank fusion, combines the two ranked lists into the result order.

Citation evidence is part of the ranking. In the word-match lane, a dataset with more published research ranks higher. The meaning lane ranks by similarity only. Download counts play no role in ranking.

Dataset summaries are different: a language model writes them from the study's own description. Every number in a summary must appear in the source record. If a summary looks wrong, please report it. If the semantic lane is unavailable, search uses keyword match only.

Some datasets say "application required" or "members only" — is the data still free?

Data TA is free to use. Dataset access follows ICPSR's rules, not ours. We show the access badge before you click. "Member only" datasets require access through an ICPSR member institution. Many "application required" datasets have some files that are open; only certain files sit behind an ICPSR application. The ICPSR study page states the exact requirements for each dataset. For use in class, we recommend sticking with "Open for everyone" or "Open for ICPSR members" datasets.

Why doesn't a dataset show citing papers — is it bad data?

No. It just means we don't have a record of a paper that cites it yet. Citation tracing is incomplete by nature, so a missing link is not evidence the dataset is unused. 8,928 of the 12,221 fully profiled studies (at least 73%) have a citing paper. Also, the most on-topic dataset for a search is often a newer, less-cited one.

Does Data TA cover the full ICPSR catalogue?

No. The graph holds 31,160 ICPSR records, and 12,221 professionally curated studies carry full profiles. The remaining records are mostly small self-published files that published papers rarely cite. Our data is a snapshot from early 2026, so newer ICPSR releases will not appear. We also answer a different question than ICPSR's own search. ICPSR search finds datasets by searching their ICPSR metadata (e.g., description, geography). Data TA finds datasets by searching the published research that used them.

Who built Data TA? How?

Jor-El Santos built the original Data TA app.

The ICPSR Knowledge Graph and its API were built by Ji Eun Kim, Taesung Ha, Jor-El Santos, and Libby Hemphill. If you have an idea for another app, please apply for access to the ICPSR Knowledge Graph API.

Something looks wrong — how do I report it or get it corrected?

Use the "Report a problem" link in the site footer, or email data-ta-project@umich.edu. Please include the search you ran and the dataset name. Counts are floors, so a higher true count is normal, not an error. A wrong citation link is an error — please report those.