← Chris Albon

An open map of congressional hearings

Visitors contribute embeddings of selected passages from the Congressional Hearings collection on GovInfo. An embedding represents a passage as 384 numbers, allowing passages with similar meanings to be compared. The model runs in the visitor’s browser; the shared collection stores the results.

What is included

We discover House, Senate, and joint committee materials across subjects in GovInfo’s CHRG collection. New and updated transcripts receive priority, alongside older material from the archive. The collection can also contain markups and errata. Discovery continues independently of visits; embedding work depends on participating visitors.

The initial selection includes identifiable congressional members’ spoken remarks and questions. It excludes witness testimony, prepared submissions, appendices, marked quotations, and passages whose speaker cannot be identified confidently. Some older documents lack usable text or speaker metadata. A completed document means its eligible passages have been processed, not that its entire text has been embedded.

GovInfo contains a selection of hearings, not every hearing held by Congress. Transcripts often become available two months to two years after the hearing. A newly indexed transcript is not necessarily a recent hearing. GovInfo’s coverage and publication guidance.

Reuse and attribution

The selected federal government remarks have a U.S. public-domain basis under 17 U.S.C. § 105. This is not a blanket license for the whole hearing: third-party material can retain copyright. We preserve speaker attribution, hearing and publication dates, document identifiers, original-source links, source hashes, and the selection method in each exported record. GovInfo’s copyright policy.

Selection is automated and conservative; it is not a legal determination for every passage. To report a misattributed passage or rights concern, email Chris with its passage or document ID.

Download and inspect

The homepage’s Download embeddings link saves the passages shown in your visit’s log. The shared JSONL export returns up to 1,000 source occurrences per response. Continue with /api/embeddings/download?after=SEQUENCE, using the final record’s sequence number, until a response is empty. Text retains its stated rights basis; no ownership of the public-domain text is claimed.

Collection progress reports discovered packages, preparation outcomes, eligible passages, and saved vectors. The overview map shows recent passages; similarity search examines up to 5,000 recent saved occurrences. Proximity indicates semantic similarity, not factual agreement or endorsement.

Downloads include the model identity and normalization method. Server checks verify assignments, source provenance, vector dimensions, and normalization. They do not independently prove that an anonymous visitor used the specified model. Earlier Wikipedia embeddings are preserved separately and are not included in this hearing collection.