US Voter Data Compilation
Aug 10, 2026
A large compilation of US voter registration data spanning multiple states (Alabama, Alaska, Arkansas, Colorado, Connecticut, and likely others) with records from 2015 through 2021. Contains detailed voter registration information including full names, residential addresses, birth year, gender, party affiliation, phone numbers, mailing addresses, precinct and district assignments, voter status, and registration dates. Data appears to have been collected from publicly available or officially obtained state voter rolls and assembled into a single archive, then shared on BreachForums (indicated by 'BF' suffix in folder name).
Data found in this dataset
Source files
Expand any file to inspect its column headers and the LLM's field-mapping reasoning, recorded during ingestion.
USVoterData_BF__data__Colorado__2017__Info__Public_EX-003-File_Layout__EX-003-Registered_Voters_List.csv26 columns1 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are middle names |
| 6 | suffix | high | [6] header 'NAME_SUFFIX', values are generational suffixes |
| 11 | address1 | high | [11] header 'HOUSE_NUM', house number in residential address |
| 12 | address2 | high | [12] header 'HOUSE_SUFFIX', part of residential address |
| 13 | address2 | high | [13] header 'PRE_DIR', directional prefix in residential address |
| 14 | address1 | high | [14] header 'STREET_NAME', part of residential address |
| 15 | address2 | high | [15] header 'STREET_TYPE', street type (e.g., St, Ave) in residential address |
| 16 | address2 | high | [16] header 'POST_DIR', directional suffix in residential address |
| 17 | address2 | high | [17] header 'UNIT_TYPE', unit type (e.g., Apt, Suite) in residential address |
| 18 | address2 | high | [18] header 'UNIT_NUM', unit number in residential address |
| 21 | city | high | [21] header 'RESIDENTIAL_CITY', city of residential address |
| 22 | state | high | [22] header 'RESIDENTIAL_STATE', 2-letter state abbreviation |
| 23 | zip | high | [23] header 'RESIDENTIAL_ZIP_CODE', 5-digit ZIP code |
| 24 | zip | high | [24] header 'RESIDENTIAL_ZIP_PLUS', ZIP+4 extension |
| 29 | dob | high | [29] header 'BIRTH_YEAR', year component of DOB |
| 31 | gender | high | [31] header 'GENDER', values are M/F or Male/Female |
| 37 | phone | high | [37] header 'PHONE_NUM', values are 10-digit phone numbers |
| 38 | address1 | high | [38] header 'MAILING_ADDRESS_1', first line of mailing address |
| 39 | address2 | high | [39] header 'MAILING_ADDRESS_2', second line of mailing address |
| 41 | city | high | [41] header 'MAILING_CITY', city of mailing address |
| 42 | state | high | [42] header 'MAILING_STATE', 2-letter state abbreviation |
| 43 | zip | high | [43] header 'MAILING_ZIP_CODE', 5-digit ZIP code |
| 44 | zip | high | [44] header 'MAILING_ZIP_PLUS', ZIP+4 extension |
| 45 | country | high | [45] header 'MAILING_COUNTRY', country name |
Notes: 47 columns total, 24 contain PII: names, addresses, DOB, gender, phone, country. All others are internal IDs, status codes, precinct info, party affiliation, dates, flags — none qualify as PII under the rules.
USVoterData_BF__data__Colorado__2017__Info__SPLIT_DISTRICTS.txt0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
Notes: No PII columns detected. File contains only jurisdictional codes and district names (Adams county, precinct codes, district types, etc.). These are administrative identifiers and not personally identifiable information per the defined field types.
USVoterData_BF__data__Colorado__2020__Voters_List__Registered_Voters_List__Part4.txt11 columns499,994 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are capitalized last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized first names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or capitalized middle names |
| 7 | fullName | high | [7] header 'VOTER_NAME', values combine LAST_NAME, FIRST_NAME, and MIDDLE_NAME |
| 19 | address1 | high | [19] header 'RESIDENTIAL_ADDRESS', values are full street addresses |
| 20 | city | high | [20] header 'RESIDENTIAL_CITY', values are city names |
| 21 | state | high | [21] header 'RESIDENTIAL_STATE', values are two-letter US state abbreviations |
| 22 | zip | high | [22] header 'RESIDENTIAL_ZIP_CODE', values are 5-digit zip codes |
| 28 | dob | high | [28] header 'BIRTH_YEAR', values are 4-digit years indicating birth year |
| 29 | gender | high | [29] header 'GENDER', values are 'Female' and 'Male' |
| 36 | phone | high | [36] header 'PHONE_NUM', values are 10-digit phone numbers |
Notes: 51 columns total, 11 contain PII: lastName, firstName, middleName, fullName, address1, city, state, zip, dob, gender, phone. All other columns are internal IDs, codes, status flags, district info, or non-PII demographic data.
USVoterData_BF__data__Colorado__2020__Voters_List__Registered_Voters_List__Part6.txt11 columns499,992 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all-caps surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are capitalized middle names |
| 7 | fullName | high | [7] header 'VOTER_NAME', values combine LAST_NAME, FIRST_NAME, and MIDDLE_NAME |
| 19 | address1 | high | [19] header 'RESIDENTIAL_ADDRESS', values are full street addresses |
| 20 | city | high | [20] header 'RESIDENTIAL_CITY', values are city names |
| 21 | state | high | [21] header 'RESIDENTIAL_STATE', all values are 'CO' (Colorado) |
| 22 | zip | high | [22] header 'RESIDENTIAL_ZIP_CODE', values are 5-digit ZIP codes |
| 28 | dob | high | [28] header 'BIRTH_YEAR', values are 4-digit years (1983-1983) |
| 29 | gender | high | [29] header 'GENDER', values are 'Female' and 'Male' |
| 36 | phone | high | [36] header 'PHONE_NUM', values are 10-digit phone numbers |
Notes: 51 columns total, 11 contain PII. All other columns are internal IDs, geographic codes, status flags, district information, or other non-PII data. No unstructured data detected.
USVoterData_BF__data__Colorado__2020__Voters_List__Registered_Voters_List__Part7.txt18 columns499,992 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are middle names |
| 6 | suffix | high | [6] header 'NAME_SUFFIX', values like 'JR' are generational suffixes |
| 7 | fullName | high | [7] header 'VOTER_NAME', values are full name strings in LAST, FIRST MIDDLE format |
| 19 | address1 | high | [19] header 'RESIDENTIAL_ADDRESS', values are street addresses |
| 20 | city | high | [20] header 'RESIDENTIAL_CITY', values are city names |
| 21 | state | high | [21] header 'RESIDENTIAL_STATE', values are state abbreviations |
| 22 | zip | high | [22] header 'RESIDENTIAL_ZIP_CODE', values are 5-digit ZIP codes |
| 28 | dob | high | [28] header 'BIRTH_YEAR', values are 4-digit birth years |
| 29 | gender | high | [29] header 'GENDER', values are 'Male'/'Female' |
| 36 | phone | high | [36] header 'PHONE_NUM', values are 10-digit phone numbers |
| 37 | address1 | medium | [37] header 'MAIL_ADDR1', mailing address line 1 (sparse but is PII address) |
| 38 | address2 | medium | [38] header 'MAIL_ADDR2', mailing address line 2 |
| 40 | city | medium | [40] header 'MAILING_CITY', mailing city |
| 41 | state | medium | [41] header 'MAILING_STATE', mailing state |
| 42 | zip | medium | [42] header 'MAILING_ZIP_CODE', mailing ZIP code |
| 44 | country | medium | [44] header 'MAILING_COUNTRY', mailing country |
Notes: 51 columns total; Colorado voter registration data. BIRTH_YEAR mapped as dob (year component only). MAIL_ADDR3 (col 39) skipped as it is a tertiary address continuation with no samples. Residential address components (HOUSE_NUM, STREET_NAME, etc., cols 11-18) are subcomponents already composed into RESIDENTIAL_ADDRESS col 19; skipped to avoid duplication. VOTER_ID, precinct/district fields, party, status, and registration dates are non-PII administrative fields and skipped.
USVoterData_BF__data__Colorado__2020__Voters_List__Registered_Voters_List__Part9.txt12 columns201,890 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all caps surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters or short names |
| 6 | suffix | high | [6] header 'NAME_SUFFIX', value 'JR' is a generational suffix |
| 7 | fullName | high | [7] header 'VOTER_NAME', values combine LAST_NAME, FIRST_NAME, MIDDLE_NAME in 'LAST, FIRST MIDDLE' format |
| 19 | address1 | high | [19] header 'RESIDENTIAL_ADDRESS', values are full street addresses |
| 20 | city | high | [20] header 'RESIDENTIAL_CITY', values are city names |
| 21 | state | high | [21] header 'RESIDENTIAL_STATE', values are two-letter US state abbreviations |
| 22 | zip | high | [22] header 'RESIDENTIAL_ZIP_CODE', values are 5-digit ZIP codes |
| 28 | dob | high | [28] header 'BIRTH_YEAR', values are 4-digit years (1923-1970) — combined with other voter data, this is birth year |
| 29 | gender | high | [29] header 'GENDER', values are 'Male'/'Female' |
| 36 | phone | high | [36] header 'PHONE_NUM', values are 10-digit US phone numbers |
Notes: 51 total columns; 9 contain PII (names, address, DOB, gender, phone). All others are voter IDs, codes, precinct info, party affiliation, status flags, and district assignments — no PII. Data is US voter registration records from Colorado (CO) with full residential addresses and birth years.
USVoterData_BF__data__Colorado__2021__Info__Public_EX-003-File_Layout__EX-003-Registered_Voters_List.csv17 columns1 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are common surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are common middle names |
| 6 | suffix | high | [6] header 'NAME_SUFFIX', values are generational suffixes (Jr, Sr, III, IV) |
| 14 | address1 | high | [14] header 'STREET_NAME', values are street names |
| 15 | address2 | high | [15] header 'STREET_TYPE', values indicate street type (St, Ave, Blvd) |
| 20 | city | high | [20] header 'RESIDENTIAL_CITY', values are city names |
| 21 | state | high | [21] header 'RESIDENTIAL_STATE', values are 2-letter US state abbreviations |
| 22 | zip | high | [22] header 'RESIDENTIAL_ZIP_CODE', values are 5-digit ZIP codes |
| 28 | dob | high | [28] header 'BIRTH_YEAR', values are 4-digit years |
| 30 | gender | high | [30] header 'GENDER', values are M/F or Male/Female |
| 36 | phone | high | [36] header 'PHONE_NUM', values are 10-digit phone numbers |
| 37 | address1 | high | [37] header 'MAILING_ADDRESS_1', values are mailing street addresses |
| 38 | address2 | high | [38] header 'MAILING_ADDRESS_2', values are mailing apartment/unit numbers |
| 40 | city | high | [40] header 'MAILING_CITY', values are mailing city names |
| 41 | state | high | [41] header 'MAILING_STATE', values are 2-letter US state abbreviations |
| 42 | zip | high | [42] header 'MAILING_ZIP_CODE', values are 5-digit ZIP codes |
Notes: 46 columns total, 16 contain PII. VOTER_ID, COUNTY_CODE, COUNTY, VOTER_NAME, STATUS_CODE, PRECINCT_NAME, ADDRESS_LIBRARY_ID, HOUSE_NUM, HOUSE_SUFFIX, PRE_DIR, POST_DIR, UNIT_TYPE, UNIT_NUM, RESIDENTIAL_ADDRESS, RESIDENTIAL_ZIP_PLUS, EFFECTIVE_DATE, REGISTRATION_DATE, STATUS, STATUS_REASON, CONFIDENTIAL, PRECINCT, SPLIT, VOTER_STATUS_ID, PARTY, PARTY_AFFILIATION_DATE, MAILING_ADDRESS_3, MAILING_ZIP_PLUS, MAILING_COUNTRY, PERMANENT_MAIL_IN_VOTER are internal IDs, timestamps, or non-PII flags and thus skipped.
USVoterData_BF__data__Connecticut__2018__Info__Town_IDs.txt0 rows
File structure
Format: FIXED·Has header: yes·Columns: 0
Notes: No PII fields detected. This file contains town/municipality reference data (ID_TOWN and NM_NAME columns) from a voter registration dataset, but the actual records are lookup tables for administrative divisions, not individual voter records with PII.
USVoterData_BF__data__Connecticut__2018__Info__town__town.txt_.txt0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
Notes: The provided data consists of town identifiers and names, with no PII fields present. Columns are purely geographic identifiers and names of towns, which do not constitute personally identifiable information under the defined categories. No emails, phone numbers, addresses, or other PII are detectable in the sample.
USVoterData_BF__data__District_of_Columbia__Washington_DC___2018__Read_Me.txt9 columns0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header 'LASTNAME', values are common surnames |
| 3 | firstName | high | [3] header 'FIRSTNAME', values are common given names |
| 4 | middleName | high | [4] header 'MIDDLE', values appear to be single characters representing initials |
| 5 | suffix | high | [5] header 'SUFFIX', values like 'Jr', 'III' match generational suffixes |
| 11 | address1 | high | [11] header 'RES STREET', values are street names/numbers |
| 12 | city | high | [12] header 'RES_CITY', values are city names |
| 13 | state | high | [13] header 'RES_STATE', values are 2-letter US state abbreviations |
| 14 | zip | high | [14] header 'RES_ZIP', values are 5-digit US ZIP codes |
| 15 | zip | high | [15] header 'RES_ZIP4', values are 4-digit ZIP+4 extensions |
Notes: 50 rows of voter registration data from US states. PII includes full names, residential addresses (street, city, state, ZIP), and generational suffixes. All other columns are internal identifiers, status codes, or voting history flags and should be skipped.
USVoterData_BF__data__Florida__2021__Voters__WAS_20210209.txt17 columns17,481 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] values are common surnames (Barnes, Studdard, Futch, Woodruff, Mitchell) |
| 3 | suffix | high | [3] value 'SR' is a generational suffix appearing after the name |
| 4 | firstName | high | [4] values are common given names (Sondra, Linda, Kristi, Judy, Dean) |
| 5 | middleName | high | [5] single letters and short names consistent with middle name/initial (K, June, Anne, L, M) |
| 7 | address1 | high | [7] values are street addresses (4195 Sunrise TRL, 4170 HWY 79, etc.) |
| 8 | address2 | medium | [8] second address line field, mostly blank consistent with apt/unit field |
| 9 | city | high | [9] values are city names (Chipley, Vernon) |
| 11 | zip | high | [11] values are 5-digit zip codes (32428, 32462) and zip+4 (324282927) |
| 12 | address1 | high | [12] mailing street address (4170 Highway 79) |
| 13 | address2 | medium | [13] mailing address line 2, no values shown but positionally consistent |
| 15 | city | high | [15] mailing city value (Vernon) |
| 16 | state | high | [16] value 'FL' is a US state abbreviation |
| 17 | zip | medium | [17] 9-digit value (324623700) consistent with mailing zip+4 code |
| 19 | gender | high | [19] values F/M are gender codes |
| 21 | skip | high | [21] values in MM/DD/YYYY format are dates of birth (01/29/1955, 03/17/1940, etc.) |
| 34 | skip | medium | [34] values (850, 321) are US telephone area codes, first part of phone number split across columns |
| 35 | skip | medium | [35] values are 7-digit numbers (2331751, 7696926) consistent with phone number suffix paired with area code in [34] |
Notes: 38 columns total, no header row (first row is data). File appears to be Florida voter registration data. Phone number is split across two columns: area code in [34] and 7-digit number in [35]. Columns [12]-[17] appear to be mailing address fields separate from residential address in [7]-[11]. Columns [20],[24]-[33] are precinct/district assignments; [23] is party affiliation; [22] is registration date; [28] is voter status — all skipped as non-PII.
USVoterData_BF__data__Florida__2021__Voting_History__ALA_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data appears to be a series of structured records without headers, containing voter registration information for Alabama (ALA). Each line includes fields such as state code, a numeric identifier, a date, a party affiliation code, and a status code. There are no explicit PII fields like names, addresses, emails, or phone numbers visible in the provided rows. The data format is consistent across rows, but without headers or identifiable PII values, no columns can be mapped to PII fields. Further inspection of additional rows or header information would be needed to accurately identify PII columns.
USVoterData_BF__data__Florida__2021__Voting_History__BAK_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII columns present. All 4 columns are non-PII: col 0 = state/county code (BAK), col 1 = internal voter ID, col 2 = election date (not a DOB), col 3 = election type code (GEN/PRI/PPP), col 4 = voted indicator (Y/N/E/A). No names, addresses, emails, phones, or other PII fields exist in this file.
USVoterData_BF__data__Florida__2021__Voting_History__BAY_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII columns detected. The 5 columns appear to be: [0] county/precinct code (e.g. 'BAY'), [1] voter registration ID (internal numeric ID), [2] election date (non-DOB timestamp), [3] election type code (GEN/PRI/OTH/PPP), [4] vote/status flag (Y/N/A/E). None of these are mappable PII fields — all are internal/administrative voter history records with no names, addresses, emails, phones, DOB, SSN, or other personal identifiers.
USVoterData_BF__data__Florida__2021__Voting_History__BRA_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This appears to be a voter participation/election history file with no PII columns. Column 0 is a county/jurisdiction code (BRA), column 1 is a numeric voter registration ID (internal), column 2 is an election date (not DOB), column 3 is election type (GEN/PRI/PPP/OTH), and column 4 is a vote participation flag (Y/N/A/E/B). No names, addresses, phones, emails, or other PII are present in this file.
USVoterData_BF__data__Florida__2021__Voting_History__BRE_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: 5-column tab-delimited voter history file with no header row. All columns are non-PII: col 0 is a constant state/source code ('BRE'), col 1 is a numeric voter registration ID, col 2 is an election date (not DOB), col 3 is an election type code (GEN/PRI/PPP/OTH), and col 4 is a participation/status code (Y/A/E/N/B). No personal identifying information is present in this file.
USVoterData_BF__data__Florida__2021__Voting_History__BRO_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII columns detected. All 5 columns are non-PII: county/jurisdiction code, voter registration ID, election date (not DOB), election type code, and voter status/activity code.
USVoterData_BF__data__Florida__2021__Voting_History__CAL_H_20210209.txt4 columns75,827 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 0 | state | high | [0] values are state abbreviations like 'CAL', consistent across rows |
| 1 | ssn | high | [1] 9-digit numeric values likely representing SSNs |
| 2 | dob | high | [2] dates in MM/DD/YYYY format represent birth dates |
| 3 | gender | high | [3] values 'M'/'F' represent gender |
Notes: 50 rows shown, likely part of a larger voter registration dataset. Columns 0-3 contain state code, SSN, date of birth, and gender. Additional columns likely exist beyond the displayed range but are not shown in this sample.
USVoterData_BF__data__Florida__2021__Voting_History__CHA_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII present. All 5 columns are non-PII: col 0 is a jurisdiction/county code ('CHA'), col 1 is a numeric voter registration ID, col 2 is an election date (not DOB), col 3 is an election type code (GEN/PRI/PPP/OTH), and col 4 is a single-letter voting status or ballot type code. This appears to be a voter history activity file with no names, contact info, or biographical data.
USVoterData_BF__data__Florida__2021__Voting_History__CIT_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file appears to be a voter history/participation log with no PII columns. All 5 columns are non-PII: a record type flag ('CIT'), an internal voter ID, an election event date (not DOB), an election type code (GEN/PRI/PPP), and a participation/ballot status code (Y/N/A/E/B). No names, addresses, emails, phones, DOB, SSN, or other mappable PII fields are present.
USVoterData_BF__data__Florida__2021__Voting_History__CLA_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: Voter history/participation file with no PII. All 5 columns are internal: county code, voter registration ID, election date, election type code, and participation/method flag. No names, addresses, emails, phones, DOB, or other mappable PII present.
USVoterData_BF__data__Florida__2021__Voting_History__CLL_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII found. The 5 columns appear to be: [0] county/jurisdiction code (CLL), [1] voter registration ID (numeric internal ID), [2] election date (non-DOB timestamp), [3] election type code (GEN/PRI/PPP/OTH), [4] vote/status code (A/E/Y/N/B). None of these columns contain PII fields — no names, addresses, phones, emails, DOB, SSN, or other personal identifiers are present in this file.
USVoterData_BF__data__Florida__2021__Voting_History__CLM_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: Voter history/activity log with 5 columns: county code, voter registration ID, election date, election type (GEN/PRI/PPP), and voting method/status code. No PII fields present — all columns are internal identifiers, event dates, and status flags.
USVoterData_BF__data__Florida__2021__Voting_History__DAD_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: Voter election history file with no PII. All 5 columns are non-PII: column 0 is a constant record-type code ('DAD'), column 1 is an internal voter ID, column 2 is an election date (not DOB), column 3 is an election type code (GEN/PRI/PPP/OTH), column 4 is a voting method or participation code (A/Y/E/B).
USVoterData_BF__data__Florida__2021__Voting_History__DES_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains no PII columns. All 5 columns are non-PII: [0] state/jurisdiction code (e.g. 'DES'), [1] internal voter registration ID, [2] election date (timestamp, not DOB), [3] election type code (GEN/PRI/PPP/OTH), [4] voting method/status code (Y/E/A/N). No names, addresses, emails, phones, DOB, SSN, or other personal identifiers are present.
USVoterData_BF__data__Florida__2021__Voting_History__DIX_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains no PII columns. The 5 columns appear to be: [0] precinct/district code (e.g. 'DIX'), [1] voter registration ID (numeric internal ID), [2] election date (non-DOB timestamp), [3] election type code (GEN/PRI/PPP/OTH), [4] voter participation/status flag (Y/N/A/E). None of these are PII fields — all are internal identifiers, election event dates, and status flags.
USVoterData_BF__data__Florida__2021__Voting_History__DUV_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: 5-column tab-delimited voter history file with no header. All columns are non-PII: col 0 = jurisdiction code (DUV), col 1 = voter registration ID, col 2 = election date (not DOB), col 3 = election type code (GEN/PRI/PPP/OTH), col 4 = participation/status flag (A/Y/E/N). No personal identifying information present in this file.
USVoterData_BF__data__Florida__2021__Voting_History__ESC_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: Voter history activity file with no PII columns. All 5 columns are internal/operational: county code (ESC), voter registration ID, election date (not DOB), election type code (GEN/PRI/PPP/OTH), and vote/status flag (Y/N/A/E/B). No names, addresses, phone numbers, emails, DOB, or other PII fields are present.
USVoterData_BF__data__Florida__2021__Voting_History__FLA_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains Florida voter history event records with no PII columns. Column 0 is a state code ('FLA'), column 1 is a numeric voter registration ID (internal identifier), column 2 is an election date (event timestamp, not DOB), column 3 is an election type code (GEN/PRI/PPP/OTH), and column 4 is a participation/method code (E/Y/A/N). No names, emails, phones, addresses, DOB, SSN, or other PII fields are present.
USVoterData_BF__data__Florida__2021__Voting_History__FRA_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a voter election history file with no PII columns. All 5 columns are non-PII: col 0 is a jurisdiction/county code (FRA), col 1 is an internal numeric voter ID, col 2 is an election date (not DOB), col 3 is an election type code (PRI/GEN/PPP/OTH), and col 4 is a participation/status code (E/Y/A/N/B). No names, addresses, emails, phones, or other personal identifiers are present.
USVoterData_BF__data__Florida__2021__Voting_History__GAD_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: 5-column tab-delimited voter history file with no header row. All columns are non-PII: col 0 is a county/jurisdiction code (GAD), col 1 is a numeric voter registration ID (internal identifier), col 2 is an election date (not DOB), col 3 is an election type code (GEN/PRI/OTH/PPP), col 4 is a participation/status code (Y/E/A/N). No names, emails, phones, addresses, or other PII are present.
USVoterData_BF__data__Florida__2021__Voting_History__GIL_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains no PII columns. Column 0 appears to be a precinct/district code (GIL), column 1 is a numeric voter ID, column 2 is an election date (non-DOB timestamp), column 3 is an election type code (GEN/PRI/PPP/OTH), and column 4 is a voted/participation flag (Y/N/A/E). None of these fields are mappable PII types.
USVoterData_BF__data__Florida__2021__Voting_History__GLA_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains no PII. All 4 columns are non-PII voter history records: col 0 is a county/jurisdiction code (GLA), col 1 is a numeric voter ID (internal identifier), col 2 is an election date (non-DOB timestamp), col 3 is an election type code (GEN/PRI/PPP), and col 4 is a voted/absentee/not-voted flag (Y/A/E/N). No names, addresses, emails, phones, DOB, SSN, or other PII fields are present.
USVoterData_BF__data__Florida__2021__Voting_History__GUL_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains no PII. It appears to be a voter history/activity log with 4 columns: [0] county/jurisdiction code (GUL), [1] internal voter ID number, [2] election date, [3] election type code (GEN/PRI/OTH/PPP), [4] ballot return/participation code (Y/N/A/E/B). None of these columns contain PII fields such as names, addresses, DOB, phone, email, SSN, or gender. All columns map to skip.
USVoterData_BF__data__Florida__2021__Voting_History__HAM_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: Voter history/participation log file with no PII columns. All 5 columns are non-PII: col 0 = county code (HAM = Hamilton), col 1 = internal voter registration ID, col 2 = election date (not DOB), col 3 = election type code (GEN/PRI/OTH/PPP), col 4 = participation status code (Y/N/E/A). No names, emails, phones, addresses, DOBs, or other mappable PII present.
USVoterData_BF__data__Florida__2021__Voting_History__HAR_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains voter activity/history records (jurisdiction code, voter ID, election date, election type, participation flag) with no PII fields mappable to the available field types. No names, emails, phones, addresses, DOB, SSN, or other personal identifiers are present.
USVoterData_BF__data__Florida__2021__Voting_History__HEN_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains no PII. The 4 columns appear to be: [0] county/jurisdiction code (e.g. 'HEN'), [1] voter registration ID (internal numeric ID), [2] election date (non-DOB date, a transactional/event timestamp), [3] election type code (GEN/PRI/PPP/OTH), [4] vote status/method code (Y/N/E/A). None of these columns contain PII fields — no names, emails, phones, addresses, DOB, SSN, gender, or other personal identifiers are present in this file. This appears to be a voter history/participation log file, not a registration file.
USVoterData_BF__data__Florida__2021__Voting_History__HER_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII found. All 5 columns are non-PII: col 0 is a constant state/county code ('HER'), col 1 is an internal voter ID, col 2 is election date (not DOB — dates cluster around known US election cycles), col 3 is election type code (GEN/PRI/PPP/OTH), col 4 is voter participation flag (Y/A/E/N). This is a voter history participation table with no names, contact info, or identifiable personal details.
USVoterData_BF__data__Florida__2021__Voting_History__HIG_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII columns detected. The 4 columns appear to be: [0] county/jurisdiction code (always 'HIG'), [1] voter registration ID (numeric internal identifier), [2] election date (non-DOB timestamp), [3] election type code (GEN/PRI/PPP/OTH), [4] vote history status code (Y/A/E/N). None of these are PII fields — all map to skip.
USVoterData_BF__data__Florida__2021__Voting_History__HIL_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: none
Notes: The provided data appears to be a tab-delimited file with no header row. The visible columns are: HIL code, a numeric voter ID, a date field, a party affiliation code (GEN, PRI, PPP, etc.), and a status flag (A, Y, E, B). None of these columns contain PII fields such as names, addresses, phone numbers, or email addresses. The data appears to be voter registration identifiers and status codes rather than personal identifying information. Therefore, no PII columns can be mapped from this dataset.
USVoterData_BF__data__Florida__2021__Voting_History__HOL_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains no PII columns. All 5 columns are non-PII voter activity/administrative data: column 0 is a county/jurisdiction code (HOL), column 1 is a voter registration ID (internal numeric ID), column 2 is an election date (non-DOB timestamp), column 3 is an election type code (GEN/PRI/PPP/OTH), and column 4 is a vote method/status code (Y/N/A/E/B). No names, addresses, phones, emails, DOB, SSN, or other PII fields are present in this file.
USVoterData_BF__data__Florida__2021__Voting_History__IND_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII present. All 5 columns are non-PII: col 0 is a constant type code ('IND'), col 1 is an internal voter registration ID, col 2 is an election date (not DOB), col 3 is an election type code (GEN/PRI/OTH/PPP), col 4 is a single-letter status/participation flag. This is a voter history transaction log with no names, addresses, contact info, or other directly identifiable fields.
USVoterData_BF__data__Florida__2021__Voting_History__JAC_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: 5-column tab-delimited voter history file with no header row. All columns are non-PII: col0=county code (JAC=Jackson), col1=voter registration ID, col2=election date (not DOB), col3=election type code (GEN/PRI/PPP/OTH), col4=voting method/status flag (Y/A/E/N). No name, address, DOB, phone, email, or other PII fields are present in this slice.
USVoterData_BF__data__Florida__2021__Voting_History__JEF_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains no PII columns. The 5 columns appear to be: [0] county/jurisdiction code (e.g. 'JEF' = Jefferson County), [1] voter registration ID (internal numeric identifier), [2] election date (a transactional/event date, not DOB), [3] election type code (GEN/PRI/PPP/OTH), [4] vote/participation status code (Y/N/A/E). None of these fields map to any PII field type — no names, emails, phones, addresses, DOB, SSN, gender, or other PII are present in this file.
USVoterData_BF__data__Florida__2021__Voting_History__LAF_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII columns detected. File appears to be a voter activity/history log with 5 columns: [0] county code (LAF), [1] voter registration ID (internal numeric ID), [2] election date (non-DOB timestamp), [3] election type code (GEN/PRI/PPP), [4] voted/ballot status flag (Y/N/E/A). None of these columns contain importable PII.
USVoterData_BF__data__Florida__2021__Voting_History__LAK_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file appears to be Alaska (LAK) voter history records with no PII columns present. The 5 columns are: jurisdiction code (LAK), internal voter registration ID, election date (non-DOB event date), election type code (GEN/PRI/PPP/OTH), and voting participation status (Y/A/E/N). None of these map to any available PII field types.
USVoterData_BF__data__Florida__2021__Voting_History__LEE_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII columns present. All 5 columns are non-PII: county name (LEE), internal voter ID, election date, election type code (PRI/GEN/OTH/PPP), and ballot/status code (A/Y/E/B). No names, emails, phones, addresses, DOBs, or other personal identifiers are present in this file.
USVoterData_BF__data__Florida__2021__Voting_History__LEO_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains voter history/participation records with no PII columns mappable to available field types. All 5 columns are non-PII: county/jurisdiction code (LEO), internal voter registration ID, election event date (not DOB), election type code (GEN/PRI/PPP), and vote participation status (Y/A/E/N).
USVoterData_BF__data__Florida__2021__Voting_History__LEV_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: Voter history file with no PII columns. All 5 columns are non-PII: column 0 is a constant type code ('LEV'), column 1 is a numeric voter ID (internal identifier), column 2 is election date (not DOB), column 3 is election type code (GEN/PRI/PPP/OTH), column 4 is participation/ballot status flag (Y/N/A/E). No names, emails, phones, addresses, DOB, SSN, or other PII present.
USVoterData_BF__data__Florida__2021__Voting_History__LIB_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file appears to be a voter history/activity log with no PII columns. Column 0 is a county/jurisdiction code (LIB), column 1 is a numeric voter ID (internal identifier), column 2 is an election date (timestamp), column 3 is an election type code (GEN/PRI/PPP/OTH), and column 4 is a voted/participation flag (Y/N/E/A). None of these columns contain directly importable PII fields (name, address, DOB, phone, email, etc.). This is likely a voter history participation file that would need to be joined with a separate voter registration file to produce PII-bearing records.
USVoterData_BF__data__Florida__2021__Voting_History__MAD_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII columns detected. This is a voter election history file with 5 columns: jurisdiction code (MAD), voter registration ID, election date, election type (GEN/PRI/PPP/OTH), and participation/ballot code (Y/E/A/N). All columns are internal identifiers, dates, and status flags — none contain personally identifiable information such as names, addresses, phone numbers, DOB, SSN, or email.
USVoterData_BF__data__Florida__2021__Voting_History__MAN_H_20210209.txt4 columns0 rows
File structure
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 0 | gender | high | [0] values are 'MAN' - clear gender marker |
| 2 | dob | high | [2] values match MM/DD/YYYY date pattern, typical voter DOB format |
| 3 | suffix | high | [3] values are party affiliation codes (DEM/REP etc), but in voter data these commonly represent suffix/polling codes - mapped as suffix per breach context |
| 4 | skip | high | [4] values are 'Y'/'A'/'E'/'N' status flags - not PII |
Notes: 50 rows shown, structure matches US voter registry format: gender, ID, DOB, party/suffix, status. Only gender, dob, and suffix (party code acting as suffix) are PII. ID column is internal voter ID (skip).
USVoterData_BF__data__Florida__2021__Voting_History__MON_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII columns present. File contains 5 tab-delimited columns: county code (col 0), voter registration ID (col 1), election date (col 2), election type code PRI/GEN/PPP/OTH (col 3), and voting participation code Y/N/A/E/B (col 4). This is a voter history/participation file, not a registration file. No names, addresses, DOB, phone, email, or other mappable PII fields are present.
USVoterData_BF__data__Florida__2021__Voting_History__MRN_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains no PII. The 5 columns are: [0] literal string 'MRN' (a record type label, constant), [1] numeric voter/member registration ID (internal identifier), [2] election dates (non-DOB timestamps), [3] election type codes (GEN/PRI/PPP/OTH — internal flags), [4] voting result/status codes (Y/N/A/E — internal flags). No names, addresses, emails, phones, DOB, SSN, or other PII fields are present.
USVoterData_BF__data__Florida__2021__Voting_History__MRT_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII columns detected. Column 0 appears to be a state/county code (MRT), column 1 is a voter registration ID, column 2 is an election date (non-DOB timestamp), column 3 is an election type code (GEN/PRI/PPP/OTH), and column 4 is a vote/participation status code (Y/A/E/N). None of these constitute mappable PII fields — they are internal identifiers, election event dates, and status flags.
USVoterData_BF__data__Florida__2021__Voting_History__NAS_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: Voter history/participation file with no PII columns. Column 0 is jurisdiction code ('NAS'), column 1 is a numeric voter registration ID (internal identifier, skip), column 2 is election date (non-DOB timestamp, skip), column 3 is election type code (GEN/PRI/PPP/OTH, skip), column 4 is ballot/vote method code (E/Y/A/N, skip). No names, addresses, contact info, or other PII present in this file.
USVoterData_BF__data__Florida__2021__Voting_History__OKA_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII columns present. This file appears to be an Oklahoma voter history log with 4 columns: [0] county/jurisdiction code (OKA), [1] internal voter ID, [2] election date (not DOB), [3] election type code (GEN/PRI/PPP/OTH), [4] vote method/status code (Y/N/E/A/B/P). None of these columns contain personal PII (names, addresses, DOB, phone, email, SSN, gender, etc.). All columns map to skip.
USVoterData_BF__data__Florida__2021__Voting_History__OKE_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
Notes: This is not a delimited file with column headers. It appears to be a series of voter records with no visible field names or consistent delimiters. All lines share the same structure: [state_code] [voter_id] [registration_date] [party_affiliation] [status_flag]. Without headers or a delimiter, we cannot map columns to PII fields. The values themselves (voter IDs, dates, party codes) do not expose direct PII without additional context. Therefore, no PII columns can be identified.
USVoterData_BF__data__Florida__2021__Voting_History__ORA_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII columns detected. The 4 columns appear to be: [0] county/jurisdiction code (ORA), [1] internal voter ID (numeric), [2] election date (timestamp, not DOB), [3] election type code (GEN/PRI/PPP/OTH), [4] ballot status code (A/E/Y/B). None of these are PII fields — they are all internal/administrative codes and non-DOB dates.
USVoterData_BF__data__Florida__2021__Voting_History__OSC_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains no PII. All 4 columns are: [0] county/jurisdiction code (e.g. 'OSC'), [1] voter registration ID (internal numeric ID), [2] election date (non-DOB timestamp), [3] election type code (GEN/PRI/PPP/OTH), [4] vote method/status code (E/Y/A/N). No names, addresses, emails, phones, DOB, SSN, or other personal identifiers are present. All columns map to skip.
USVoterData_BF__data__Florida__2021__Voting_History__PAL_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file appears to be a voter history/participation log, not a voter registration file. It contains no PII columns — only: col[0] = state/jurisdiction code ('PAL'), col[1] = internal voter ID (numeric, skip), col[2] = election date (timestamp, skip), col[3] = election type code (GEN/PRI/PPP/OTH, skip), col[4] = participation/ballot status code (E/Y/A, skip). No names, addresses, emails, phones, DOB, gender, or any other PII fields are present in this file.
USVoterData_BF__data__Florida__2021__Voting_History__PAS_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII present. All 5 columns are non-PII: col 0 is a constant state/file code ('PAS'), col 1 is a numeric voter registration ID, col 2 is an election date (not DOB), col 3 is an election type code (GEN/PRI/PPP/OTH), col 4 is a participation/status flag (Y/N/A/E). This is a voter history participation file with no names, contact info, addresses, or demographic PII.
USVoterData_BF__data__Florida__2021__Voting_History__PIN_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: Voting history file with no PII columns. Column 0 is a literal 'PIN' label (constant string, not a field name), column 1 is internal voter ID, column 2 is election date (not DOB), column 3 is election type code (GEN/PRI/OTH/PPP), column 4 is vote/ballot status code (Y/A/E/N/B/P). No names, addresses, emails, phones, or other PII present.
USVoterData_BF__data__Florida__2021__Voting_History__POL_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII columns present. This file contains voter participation history records: a record-type constant ('POL'), internal voter ID numbers, election dates (not DOBs), election type codes (GEN/PRI/PPP/OTH), and voting status codes (A/Y/E/B). No names, contact info, addresses, or other personal identifiers are present in this file.
USVoterData_BF__data__Florida__2021__Voting_History__PUT_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII present. File contains voter history records: column 0 is a state/record-type code ('PUT'), column 1 is an internal voter registration ID, column 2 is an election date (not DOB), column 3 is election type code (GEN/PRI/PPP/OTH), and column 4 is a voting method or status flag (Y/E/A/N/B/P). All columns are skip.
USVoterData_BF__data__Florida__2021__Voting_History__SAN_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: Voter election history file with no PII. All 5 columns are non-PII: county code (col 0), voter registration ID (col 1), election date/event date (col 2, not DOB), election type code (col 3), and ballot/status code (col 4). No names, addresses, phones, emails, or other searchable PII present.
USVoterData_BF__data__Florida__2021__Voting_History__SAR_H_20210209.txt0 rows
File structure
Notes: This file is unstructured free-form text with no consistent column delimiter. Despite containing voter record-like patterns (SAR prefixes, numeric IDs, dates, codes), the data lacks a fixed columnar structure across lines. The sample shows free-text lines with varying field counts and positions, confirming it's not a delimited CSV/TSV. No PII columns can be mapped due to absence of header or consistent layout.
USVoterData_BF__data__Florida__2021__Voting_History__SEM_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII found. All 4 columns are internal/administrative voter history data: col 0 = county/jurisdiction code (SEM), col 1 = voter registration ID (numeric internal identifier), col 2 = election date (non-DOB timestamp), col 3 = election type code (GEN/PRI/OTH/PPP), col 4 = vote method/status code (Y/N/A/E/B). None of these fields contain personal PII such as names, addresses, DOB, phone, email, or SSN.
USVoterData_BF__data__Florida__2021__Voting_History__STJ_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains no PII. All 4 columns are non-PII: [0] county/jurisdiction code (STJ), [1] internal voter ID number, [2] election date (timestamp, not DOB), [3] election type code (GEN/PRI/PPP/OTH), [4] voting method/status code (E/A/Y/N). No names, addresses, phones, emails, DOB, or other personal identifiers are present in this file.
USVoterData_BF__data__Florida__2021__Voting_History__STL_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII columns detected. This file appears to be a voter history/activity log containing: [0] county/jurisdiction code (STL), [1] voter registration ID (internal numeric ID), [2] election date (non-DOB timestamp), [3] election type code (GEN/PRI/PPP/OTH), [4] participation/method code (Y/A/E/N/B). None of these columns contain mappable PII fields — no names, emails, phones, addresses, DOB, SSN, or gender. All columns are internal identifiers, dates, or status codes and should be skipped.
USVoterData_BF__data__Florida__2021__Voting_History__SUM_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: Voter history/participation log file. All 5 columns are non-PII: col 0 is a record type code ('SUM'), col 1 is internal voter ID, col 2 is election date (not DOB), col 3 is election type code (GEN/PRI/PPP/OTH), col 4 is participation status code (Y/N/E/A). No names, contact info, or other PII present.
USVoterData_BF__data__Florida__2021__Voting_History__SUW_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: Voter history/participation file with no header. All 5 columns are non-PII: county code, voter registration ID, election date, election type code (GEN/PRI/PPP/OTH), and participation method code (Y/E/A/N). No names, contact info, addresses, or other importable PII fields present.
USVoterData_BF__data__Florida__2021__Voting_History__TAY_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains no PII columns. Column 0 appears to be a county/jurisdiction code ('TAY' = Taylor County), column 1 is an internal voter registration ID (skip), column 2 is an election date (non-DOB timestamp, skip), column 3 is an election type code (GEN/PRI/PPP, skip), and column 4 is a voted/status flag (Y/A/E/N, skip). No names, addresses, emails, phones, DOB, SSN, or other personal PII are present in this file — it is a voter history/participation log, not a voter registration record.
USVoterData_BF__data__Florida__2021__Voting_History__UNI_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a voter election history/participation file with no PII columns. All 5 columns are non-PII: col 0 = jurisdiction code ('UNI'), col 1 = numeric voter ID (internal identifier), col 2 = election date (non-DOB event date), col 3 = election type code (GEN/PRI/PPP), col 4 = participation/ballot status code (Y/N/E/A). No names, addresses, emails, phones, DOB, SSN, or other personal identifiers are present.
USVoterData_BF__data__Florida__2021__Voting_History__VOL_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: 5-column tab-delimited voter history file with no header. Columns contain: record type code ('VOL'), numeric voter ID, election date, election type code (GEN/PRI/PPP/OTH), and voting method/status code (A/Y/E/N/B). No PII columns present in this file — no names, addresses, DOB, phone, email, or other personal identifiers.
USVoterData_BF__data__Florida__2021__Voting_History__WAK_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains no PII columns. The 5 columns appear to be: [0] county/jurisdiction code (WAK), [1] voter registration ID (numeric internal identifier), [2] election date (timestamp, not DOB), [3] election type code (GEN/PRI/PPP/OTH), [4] vote method/status code (Y/E/A/N). None of these are PII fields — all map to skip.
USVoterData_BF__data__Florida__2021__Voting_History__WAL_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII columns detected. All 5 columns are non-PII: county/jurisdiction code, voter registration ID, election date (not a DOB), election type code, and vote method/status code. This appears to be a voter history/participation log file rather than a voter registration file — it records which elections a voter participated in, not their personal details.
USVoterData_BF__data__Florida__2021__Voting_History__WAS_H_20210209.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a voter participation history file with no PII columns. All 5 columns are non-PII: column 0 is a jurisdiction code ('WAS'), column 1 is a voter registration ID (internal numeric identifier), column 2 is an election date (not DOB), column 3 is an election type code (PRI/GEN/PPP/OTH), and column 4 is a participation/result flag (Y/A/E/N). No names, addresses, emails, phones, or other PII are present.
USVoterData_BF__data__Georgia__2018__GeorgiaVoter.txt9 columns0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are capitalized last names (BOSWELL, CARLAN, CAGLE, COCHRAN, DORSEY) |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names (J, SWAYNE, RONNIE, LINTON, HENRIETTA) |
| 5 | middleName | high | [5] header 'MIDDLE_MAIDEN_NAME', values are single capital letters or full middle names (R, G, D, A, GIBSON) |
| 12 | city | high | [12] header 'RESIDENCE_CITY', values are city names (MAYSVILLE) |
| 13 | zip | high | [13] header 'RESIDENCE_ZIPCODE', values are 9-digit ZIP+4 format (305581759, 305585009, etc.) |
| 14 | dob | high | [14] header 'BIRTHDATE', values are 4-digit years (1937, 1940, 1954, 1952, 1945) |
| 17 | gender | high | [17] header 'GENDER', values are single-letter codes (M, M, M, M, F) |
| 57 | city | high | [57] header 'MAIL_CITY', values are city names (MAYSVILLE, LAWRENCEVILLE) |
| 59 | zip | high | [59] header 'MAIL_ZIPCODE', values are 5-digit or 9-digit ZIP codes (305580012, 300437176) |
Notes: 63 total columns. PII identified in 9 columns: lastName, firstName, middleName, city (residence and mailing), zip (residence and mailing), dob (birth year), gender. All other columns are internal codes, district IDs, dates (non-DOB), race descriptors, status flags, or empty fields — mapped to skip.
USVoterData_BF__data__Kansas__2018__Info__keys__Standard.csv17 columns0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 7 | firstName | high | Header 'text_name_first' maps to firstName, contains given names |
| 8 | middleName | high | Header 'text_name_middle' maps to middleName |
| 9 | lastName | high | Header 'text_name_last' maps to lastName |
| 10 | suffix | high | Header 'cde_name_suffix' maps to suffix, contains generational suffixes like Jr, II, III |
| 12 | gender | high | Header 'cde_gender' maps to gender, contains M/F/U values |
| 28 | address1 | high | Header 'text_res_address_nbr' maps to address1, contains residential address numbers |
| 31 | address1 | high | Header 'text_street_name' maps to address1, contains street names |
| 36 | city | high | Header 'text_res_city' maps to city |
| 37 | state | high | Header 'cde_res_state' maps to state |
| 38 | zip | high | Header 'text_res_zip5' maps to zip |
| 45 | city | high | Header 'text_mail_city' maps to city |
| 46 | state | high | Header 'cde_mail_state' maps to state |
| 47 | zip | high | Header 'text_mail_zip5' maps to zip |
| 62 | skip | high | Header 'date_of_birth' maps to dob |
| 64 | skip | high | Header 'text_phone_area_code' combined with exchange/last four forms complete phone number |
| 65 | skip | high | Header 'text_phone_exchange' combined with area code/last four forms complete phone number |
| 66 | skip | high | Header 'text_phone_last_four' combined with area code/exchange forms complete phone number |
Notes: 50 rows shown; this is voter registration data with PII columns mapped. Additional rows likely follow the same pattern. Non-PII columns like internal IDs, election codes, and district information are skipped.
USVoterData_BF__data__Kansas__2018__Kansas.txt22 columns1,822,596 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | firstName | high | [3] header 'text_name_first', values are common given names like Craig, George, Georgia |
| 4 | middleName | high | [4] header 'text_name_middle', values are middle names/initials like Melton, Edward, E, L |
| 5 | lastName | high | [5] header 'text_name_last', values are surnames like Abbott, Achor |
| 6 | suffix | medium | [6] header 'cde_name_suffix', generational/credential suffix field following name columns |
| 7 | gender | high | [7] header 'cde_gender', values are M/F gender codes |
| 9 | address1 | high | [9] header 'text_res_address_nbr', street number component of residential address; combined with street name fields forms address1 |
| 11 | address1 | medium | [11] header 'cde_street_dir_prefix', directional prefix (e.g. N) part of residential street address |
| 12 | address1 | high | [12] header 'text_street_name', street name component of residential address like Fairway, East, 3rd |
| 13 | address1 | high | [13] header 'cde_street_type', street type suffix (Ave, St) part of residential address |
| 16 | address2 | medium | [16] header 'text_res_unit_nbr', unit/apartment number component of residential address |
| 17 | city | high | [17] header 'text_res_city', values are city names like Iola, Chanute |
| 18 | state | high | [18] header 'cde_res_state', values are US state abbreviations like KS |
| 19 | zip | high | [19] header 'text_res_zip5', values are 5-digit ZIP codes like 66749, 66720 |
| 23 | address1 | high | [23] header 'text_mail_address1', mailing address line 1 values like PO Box 165 |
| 24 | address2 | medium | [24] header 'text_mail_address2', mailing address line 2 |
| 27 | city | high | [27] header 'text_mail_city', mailing city values like Iola |
| 28 | state | high | [28] header 'cde_mail_state', mailing state abbreviations like KS |
| 29 | zip | high | [29] header 'text_mail_zip5', mailing ZIP codes like 66749 |
| 37 | skip | high | [37] header 'date_of_birth', values are dates in MM/DD/YYYY format like 07/09/1948 |
| 39 | skip | high | [39] header 'text_phone_area_code', area code component of phone number like 620 |
| 40 | skip | high | [40] header 'text_phone_exchange', exchange component of phone number like 365 |
| 41 | skip | high | [41] header 'text_phone_last_four', last four digits of phone number like 1600, 5920 |
Notes: 72 columns total. Phone number is split across three columns (area code [39], exchange [40], last four [41]). Residential address is split across house number [9], direction prefix [11], street name [12], street type [13], unit [16], city [17], state [18], zip [19]. Mailing address has its own set of columns [23-29]. Polling place address columns [56-62] are institutional/public locations, not personal PII, and are skipped. District and election history columns are voter registration metadata, skipped. Column [0] 'db_logid' contains county names not personal IDs. Column [2] 'cde_name_title' is an honorific/salutation prefix field (Mr/Mrs/Ms), skipped. Column [8] 'text_registrant_id' is an internal voter registration ID, skipped.
USVoterData_BF__data__Maryland__2018__Info__Voter_Registry.txt23 columns0 rows
File structure
Format: CSV·Delimiter: ---·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 1 | lastName | high | [1] header 'LastName', voter registration data contains surnames |
| 2 | firstName | high | [2] header 'FirstName', voter registration data contains given names |
| 3 | middleName | high | [3] header 'MiddleName', voter registration data contains middle names |
| 4 | suffix | high | [4] header 'Suffix', generational/credential suffixes (Jr, Sr, III, etc.) |
| 5 | address1 | high | [5] header 'HouseNumber' — component of street address; combined with street name/type for full address1 |
| 6 | address1 | high | [6] header 'HouseSuffix' — component of street address (e.g., A, B, 1/2) |
| 7 | address1 | high | [7] header 'StreetPreDirection' — component of street address (N, S, E, W) |
| 8 | address1 | high | [8] header 'StreetName' — component of street address |
| 9 | address1 | high | [9] header 'StreetType' — component of street address (St, Ave, Blvd, etc.) |
| 10 | address1 | high | [10] header 'StreetPostDirection' — component of street address (N, S, E, W) |
| 11 | address2 | high | [11] header 'UnitType' — component of secondary address (Apt, Suite, Unit) |
| 12 | address2 | high | [12] header 'UnitNumber' — component of secondary address |
| 13 | address1 | high | [13] header 'NonStandardAddress' — fallback address field for non-standard formats |
| 14 | city | high | [14] header 'ResidentialCity', voter registration residential city |
| 15 | state | high | [15] header 'ResidentialState', 2-letter state code for residential address |
| 16 | zip | high | [16] header 'ResidentialZip', 5-digit postal code for residential address |
| 17 | zip | high | [17] header 'ResidentialZipPlus', +4 extension of residential zip code |
| 18 | address1 | high | [18] header 'MAILINGADDRESS', full mailing address as single field |
| 19 | city | high | [19] header 'MAILINGCITY', mailing city |
| 20 | state | high | [20] header 'MAILINGSTATE', 2-letter state code for mailing address |
| 21 | zip | high | [21] header 'MAILINGZIP', 5-digit postal code for mailing address |
| 22 | zip | high | [22] header 'MAILINGZIPPLUS', +4 extension of mailing zip code |
| 25 | gender | high | [25] header 'Gender', voter registration gender field (M/F or similar codes) |
Notes: US voter registration data compilation. 39 columns total; 22 mapped to PII (primarily name, address, and gender components). Columns 0 (VTR_ID), 23 (StatusCode), 24 (Party), 26-34 (Congressional/Legislative/district codes), 35-36 (registration dates), 37-39 (election/voting history records), and 38 (County) are skipped as internal identifiers, administrative codes, or non-PII registration metadata. Address fields (5-13, 18, 19-22) represent components and variations of residential and mailing addresses; extraction logic should reconstruct full address1 and address2 where applicable.
USVoterData_BF__data__Maryland__2018__Info__Voting_History.txt0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
Notes: All columns represent internal identifiers, election metadata, and administrative codes — no PII fields present. VoterId is a unique internal ID, ElectionDate/Date of Voting are timestamps, PoliticalParty/ElectionType/VotingMethod are categorical codes, and location fields (Precinct/JurisdictionCode/CountyName) refer to geographic entities, not personal addresses. No emails, phone numbers, names, DOB, or other PII values detected in the sample.
USVoterData_BF__data__Maryland__2018__Maryland.txt11 columns0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 1 | lastName | high | [1] header 'LastName', values are capitalized surnames |
| 2 | firstName | high | [2] header 'FirstName', values are capitalized given names |
| 3 | middleName | high | [3] header 'MiddleName', values are capitalized middle names |
| 4 | suffix | high | [4] header 'Suffix', values are generational suffixes like 'JR' |
| 14 | city | high | [14] header 'ResidentialCity', values are city names |
| 15 | state | high | [15] header 'ResidentialState', values are two-letter state abbreviations |
| 16 | zip | high | [16] header 'ResidentialZip', values are 5-digit ZIP codes |
| 19 | city | high | [19] header 'MAILINGCITY', values are city names |
| 20 | state | high | [20] header 'MAILINGSTATE', values are two-letter state abbreviations |
| 21 | zip | high | [21] header 'MAILINGZIP', values are 5-digit ZIP codes |
| 25 | gender | high | [25] header 'Gender', values are 'Male'/'Female' |
Notes: 41 total columns, 11 contain PII: full names, residential addresses, mailing addresses, gender. All other columns are internal IDs, election codes, geographic codes, or timestamps and are skipped per rules.
USVoterData_BF__data__Maryland__2018__Voting_History__Maryland_History.txt0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
Notes: All columns represent internal IDs, election metadata, jurisdiction codes, and administrative classifications. No PII fields (names, addresses, dates of birth, contact info) are present in the first 50 rows. Columns like Voter ID are internal numeric identifiers, not usernames or personal identifiers. Election dates, precinct codes, and county names are administrative and not personally identifying on their own.
USVoterData_BF__data__Michigan__2015__michigan.txt13 columns7,408,330 rows
File structure
Format: FIXED·Has header: no·Columns: 14
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 0 | lastName | high | [0-34] left-aligned space-padded surnames: SCHOUWINK, SNYDER, IRVIN, etc. |
| 1 | firstName | high | [34-54] left-aligned space-padded given names: KAREN, CHARLOTTE, YVONNE, etc. |
| 2 | middleName | high | [54-75] middle names or initials: ANN, N, FRANCES, MADISON, blank |
| 3 | dob | medium | [75-79] 4-digit birth year only: 1959, 1931, 1960, etc. — partial DOB (year only) |
| 4 | gender | high | [79-80] single character M/F gender code |
| 5 | address1 | high | [88-127] house number [88-97] + street name [97-127] together form address1 street line |
| 6 | skip | high | [127-148] street type suffix (RUN, ST, AVE, PKWY, RD) — part of address, merged into address1 |
| 7 | city | high | [148-183] city names: ALLEGAN, DETROIT, TRAVERSE CITY, NEW BALTIMORE, etc. |
| 8 | state | high | [183-185] two-letter US state abbreviation: MI throughout |
| 9 | zip | high | [185-190] 5-digit ZIP codes: 49010, 48238, 49328, etc. |
| 10 | address2 | medium | [190-220] mailing/foreign street address when present: LANGGATAN 12 (row 9), otherwise blank |
| 11 | address2 | medium | [220-255] mailing city/region when present: 27143 YSTAD (row 9), otherwise blank |
| 12 | country | medium | [255-275] country field when present: SWEDEN (row 9), otherwise blank (implied USA) |
Notes: Michigan voter registration fixed-width file. Birth year only (no full DOB). Address1 spans house number [88-97] + street name [97-127] + street type [127-148]. Rows 80-88 appear to be a registration date (MMDDYYYY) — non-PII voter admin field, skipped. Columns 275+ contain voter IDs, precinct/district codes, registration status — all skipped as non-PII. Row 9 (RISINGER) has a foreign mailing address (Sweden) populating address2/country fields that are blank for all other rows.
USVoterData_BF__data__Michigan__2017__Info__List_Of_All_Counties_And_Codes.txt0 rows
File structure
Format: FIXED·Has header: no·Columns: 0
Notes: No PII fields detected. This section contains Michigan county names and codes (lookup/reference data), not personally identifiable information. All columns are non-PII and should be skipped.
USVoterData_BF__data__Michigan__2017__Info__List_Of_All_Jurisdictions_And_Codes.txt0 rows
File structure
Format: FIXED·Has header: no·Columns: 0
Notes: No PII fields detected. Data contains only township/precinct names and numeric codes. This appears to be a geographic reference table or precinct listing, not voter records with personal information. All columns map to skip (non-PII administrative/reference data).
USVoterData_BF__data__Michigan__2017__Info__List_Of_All_SchoolDistricts_And_Codes.txt0 rows
File structure
Format: FIXED·Has header: no·Columns: 0
Notes: No PII fields detected. This file contains school district names and numeric codes (likely district IDs or administrative codes), not personal information. Despite breach context indicating voter registration data, this particular file segment appears to be a reference lookup table or administrative directory of Michigan school districts. All columns are non-PII and should be skipped.
USVoterData_BF__data__Michigan__2017__Michigan.txt15 columns7,359,197 rows
File structure
Format: FIXED·Has header: no·Columns: 16
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 0 | lastName | high | [0-35] Left-aligned space-padded surnames: SCHOUWINK, IRVIN, VANDERVERE, WRIGHT, etc. |
| 1 | firstName | high | [35-55] Left-aligned first names: KAREN, YVONNE, STEWART, MARGO, etc. |
| 2 | middleName | high | [55-75] Middle names/initials: M, N, A, ANN, ANNE, LEE, RALPH, etc., space-padded |
| 3 | dob | high | [75-79] 4-digit birth years: 1959, 1960, 1942, 1939, 1942, etc. (birth year only) |
| 4 | gender | high | [79-80] Single char M/F gender codes |
| 5 | skip | high | [80-88] MMDDYYYY registration dates — not a PII field type |
| 6 | address1 | high | [88-145] House number + street name + street type combine to form address1: e.g. 1843 FAWNBROOKE RUN, 15744 PINEHURST ST |
| 7 | skip | high | [125-145] Street type suffix (RUN, ST, AVE, RD) — part of address1, subsumed above |
| 8 | city | high | [145-185] City names: ALLEGAN, DETROIT, TRAVERSE CITY, NEW BALTIMORE, etc., space-padded to 40 chars |
| 9 | state | high | [185-187] 2-char US state abbreviations: MI throughout |
| 10 | zip | high | [187-192] 5-digit ZIP codes: 49010, 48238, 49686, 48202, etc. |
| 11 | address2 | medium | [192-252] Mailing/alternate address line present on some records: LANGGATAN 12, PO BOX 263, blank on most |
| 12 | city | medium | [252-292] Mailing city+state field: 27143 YSTAD, HESSEL MI — present only when address2 populated |
| 13 | zip | medium | [292-297] Mailing ZIP: 49745 on line 12, blank most records |
| 14 | country | medium | [297-307] Country field: SWEDEN on line 6, blank for domestic records |
Notes: Michigan voter registration fixed-width file. Birth year only (not full DOB) at pos 75-79. Address1 spans house number (88-95) + street name (95-125) + street type (125-145). Some records have mailing/overseas addresses in cols 192-307 (lines 6 and 12 show SWEDEN and PO BOX examples). Registration date at 80-88 is not a PII field type. District/precinct/status codes follow pos 307 and are all skipped. Street type column (125-145) is part of address1 and merged accordingly.
USVoterData_BF__data__Michigan__2017__Voting_History__History.txt0 rows
File structure
Format: FIXED·Has header: no·Columns: 0
Notes: No PII fields detected. Data appears to be internal identifiers (voter registration IDs, precinct/district codes, status flags) with no extractable personally identifiable information in the visible columns. The breach context indicates voter registration data exists, but these particular columns contain only numeric codes and administrative flags, not names, addresses, phone numbers, DOB, or other PII.
USVoterData_BF__data__Michigan__2019__EntireStateVoterHistory.csv0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
Notes: No PII fields detected in the provided rows. All columns appear to be internal IDs, codes, dates, and flags related to voter registration records. None of the columns contain personal identifiable information such as names, addresses, dates of birth, phone numbers, email addresses, SSNs, passwords, usernames, or other PII categories.
USVoterData_BF__data__Michigan__2019__EntireStateVoters.csv9 columns0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 0 | lastName | high | [0] header 'LAST_NAME', values are common surnames |
| 1 | firstName | high | [1] header 'FIRST_NAME', values are common given names |
| 2 | middleName | high | [2] header 'MIDDLE_NAME', values are middle names |
| 3 | suffix | high | [3] header 'NAME_SUFFIX', values are generational suffixes (JR) |
| 4 | dob | high | [4] header 'YEAR_OF_BIRTH', values are years (1962-1989) |
| 5 | gender | high | [5] header 'GENDER', values are M/F |
| 15 | city | high | [15] header 'CITY', values are city names |
| 16 | state | high | [16] header 'STATE', values are state abbreviations (MI) |
| 17 | zip | high | [17] header 'ZIP_CODE', values are 5-digit zip codes |
Notes: 48 total columns, 7 contain PII: lastName, firstName, middleName, suffix, dob, gender, city, state, zip. All others are district codes, precinct numbers, voter status flags, and identifiers (non-PII).
USVoterData_BF__data__Michigan__2019__Info__countycd.lst.csv0 rows
File structure
Notes: This is a list of county names and codes, not a structured dataset with PII. The file contains only geographic identifiers (county names and numeric codes) with no personal information such as names, addresses, or contact details. There are no columns to map to PII fields.
USVoterData_BF__data__Michigan__2019__Info__electionscd.lst.csv0 rows
USVoterData_BF__data__Michigan__2019__Info__jurisdcd.lst.csv0 rows
File structure
Format: FIXED·Has header: no·Columns: 0
Notes: No PII fields detected. Data contains only internal administrative codes (positions 0-7) and township/jurisdiction names (positions 7-43). This appears to be a lookup table or reference file mapping precinct/district codes to geographic locations, not personal information records. All columns are non-PII and should be skipped.
USVoterData_BF__data__Michigan__2019__Info__schoolcd.lst.csv0 rows
File structure
Format: FIXED·Has header: no·Columns: 0
Notes: No PII fields detected. Data contains only school district identifiers and names. This appears to be an administrative/institutional dataset rather than voter registration records with personal information. All columns are non-PII (internal IDs and organization names) and should be skipped.
USVoterData_BF__data__Michigan__2019__Info__villagecd.lst.csv2 columns0 rows
File structure
Format: FIXED·Has header: no·Columns: 2
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 0 | skip | high | [0-20] numeric voter ID or record sequence number |
| 1 | city | medium | [20-76] space-padded city names (ADDISON, AKRON, ALANSON, ALLEN, ALMONT, ALPHA, APPLEGATE, ARMADA, ASHLEY, ATHENS, AUGUSTA) consistent with voter registration data |
Notes: Only 2 columns visible in this sample; likely truncated excerpt. Expected full voter record should include firstName, lastName, address1, address2, state, zip, phone, dob, gender. City column is PII-adjacent (part of address); included per breach context (residential address data present). Numeric prefix appears to be voter ID or sequence identifier.
USVoterData_BF__data__Missouri__2017__StateWide_VotersList_02262017_2-48-56_AM.txt8 columns0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | firstName | high | [2] header 'First Name', values are capitalized names |
| 3 | middleName | high | [3] header 'Middle Name', values are single letters or names |
| 4 | lastName | high | [4] header 'Last Name', values are capitalized surnames |
| 5 | suffix | high | [5] header 'Suffix', values appear to be generational suffixes (though currently empty in sample) |
| 15 | city | high | [15] header 'Residential City', values are city names |
| 16 | state | high | [16] header 'Residential State', values are two-letter state abbreviations |
| 17 | zip | high | [17] header 'Residential ZipCode', values are 5-digit ZIP codes |
| 22 | skip | high | [22] header 'Birthdate', values are in MM/DD/YYYY date format |
Notes: 53 columns total, 8 contain PII. Columns 0 (County), 1 (Voter ID), 6-21 (address components), 23-52 (voter history dates) are skip per rules (internal IDs, timestamps, or non-PII address fragments).
USVoterData_BF__data__Nevada__2015__Nevada_Voters.txt13 columns1,160,837 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 12 | firstName | high | [12] header 'FIRST_NAME', values are names |
| 13 | middleName | high | [13] header 'MIDDLE_NAME', values are names |
| 14 | lastName | high | [14] header 'LAST_NAME', values are names |
| 16 | gender | high | [16] header 'SEX', values are 'M' and presumably 'F' elsewhere |
| 18 | dob | high | [18] header 'BIRTH_YEAR', values are 4-digit years |
| 19 | phone | high | [19] header 'PHONE_NUM', values are 10-digit numbers with area code |
| 20 | address1 | high | [20] header 'RES_STREET_NUM', values are street numbers |
| 22 | address1 | high | [22] header 'RES_STREET_NAME', values are street names |
| 23 | address2 | high | [23] header 'RES_ADDRESS_TYPE', values are street types (ST, LN, AVE, etc.) |
| 24 | address2 | high | [24] header 'RES_UNIT', values are apartment/unit numbers |
| 25 | city | high | [25] header 'RES_CITY', values are city names |
| 26 | state | high | [26] header 'RES_STATE', values are state abbreviations |
| 27 | zip | high | [27] header 'RES_ZIP_CODE', values are 5-digit zip codes |
Notes: 80 columns total, 12 contain PII. Columns 0-11, 15, 17-79 are internal IDs, election codes, flags, timestamps, or other non-PII data. The 'NAME_SUFFIX' column is present but appears empty or mislabeled in this sample; if populated with generational suffixes (Jr, Sr, III, etc.), it would map to 'suffix'.
USVoterData_BF__data__Nevada__2018__Campaign_Finances__CampaignFinance.Cnddt.csv2 columns0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 1 | firstName | high | [1] values are common given names (Michael, Richard, Carlo...), header 'Michael' suggests first name |
| 2 | lastName | high | [2] values are common surnames (Douglas, Ziser, Poliark...), header 'Douglas' suggests last name |
USVoterData_BF__data__Nevada__2018__Campaign_Finances__CampaignFinance.Cntrbt.csv0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
Notes: The provided data appears to be a structured CSV file, but all visible columns represent transaction IDs, monetary amounts, contribution types, and numerical codes. No identifiable PII fields (email, phone, name, address, etc.) are present in the sample rows. The data resembles financial contribution records without personal identifiers. All columns fall under skip criteria (transactional IDs, amounts, internal codes).
USVoterData_BF__data__Nevada__2018__Campaign_Finances__CampaignFinance.Cntrbtrs.csv3 columns0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 0 | firstName | high | [0] values are common given names (Bonnie, John, Robert, James, Howard, Jeromi, Robert, Dave, Chris, April, Marilee, Fred, Ernest, Edward, James, Russell, Robert, Alan, Jeffrey, F.A, Mike, Mendy, Don, David, Dean, Charles, Daniel, Mark, Kathleen, Alfred, James, Curtis, Howard, Shirley, E. Alan, Michael, John, Natalie, David, John, James, Ron, Sheila, Marcus, Mark, Michelle, James, Tom, Mary, Courtney, Jim, Stephen, Ken, Gene, Jack, James, Jack, Joe, Dave, Ronald, Dina, Jack, Carl, Toni, Peter, Warren, John, Neal, Charles, Paul, Janell, Jerry, David, john, Camile, jenny) |
| 1 | middleName | high | [1] contains middle initials and full middle names (B, C, E, H, S, D, A, P, D, K, A, L, F, G, S, R, B, F) |
| 2 | lastName | high | [2] contains valid last names (Jacobs, Mueller, Paganelli, Tiras, Clark, Elias, Offerdahl, Hengst, Hubbard, Wheeler, Marriner, Plastiras, Lucking, Kovacs, Smith, Grossman, Leutheuser, Giudici, Bishop, Cashell, Ferris, Gussow, Kuckhoff, Chamberlain, Elliott, Kornstein, Thompson, Meiling, Otto, Wong, Sonnenshein, Fincham, Sperry, Jervey, Mclaclan, Marguleas, Dale, Tiras, Colarchik, Richard, Tiras, Menath, Houston, Routsis, Salvucci, Cauley, Salvucci, Sheehan, Adams, Mollett, Neumann, Ronsheimer, Hara, Cobb, Ragusa, Yohey, Detrick, Anderson, Houston, calvert, Lewis/Richards Jewelers, hubach) |
USVoterData_BF__data__Nevada__2018__Campaign_Finances__CampaignFinance.Expn.csv0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
Notes: The provided data appears to be a financial ledger or expense report, containing transaction IDs, monetary amounts, and expense types. There are no identifiable PII fields (names, addresses, emails, phone numbers, etc.) present in the visible rows. All columns contain internal IDs, transaction dates, monetary values, and expense classifications.
USVoterData_BF__data__Nevada__2018__Campaign_Finances__CampaignFinance.Grp.csv3 columns1,135 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | fullName | high | [2] values are full names like 'Shirlanda Walker', 'James L. Wadhams', etc. |
| 3 | gender | high | [3] values are single letters 'T'/'F' representing gender |
| 4 | city | high | [4] values are city names like 'Northbrook', 'Sacramento', etc. |
Notes: 50 rows shown, 3 PII columns identified. Column 0 appears to be an ID number, columns 1 contains PAC names which are organizational entities and not personal PII, and should be skipped. Gender is represented as T/F codes.
USVoterData_BF__data__Nevada__2018__Campaign_Finances__CampaignFinance.Rpr.csv0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
Notes: The provided data appears to be a structured CSV file without headers. The columns do not contain any identifiable PII fields such as names, addresses, email addresses, phone numbers, dates of birth, SSNs, passwords, usernames, gender, suffix, or Facebook IDs. The data seems to consist of numerical identifiers, report types, years, dates, and flags, which are not classified as PII according to the given criteria. Therefore, no PII columns can be mapped.
USVoterData_BF__data__Nevada__2018__Nevada.csv11 columns1,743,936 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 1 | lastName | high | [1] header 'Clark', values are consistent surnames |
| 2 | firstName | high | [2] header 'TARIK', values are common given names |
| 3 | middleName | high | [3] header 'A', values are single letters used as middle initials |
| 4 | suffix | medium | [4] header 'ABI-KARAM', values appear to be generational suffixes or credentials (e.g., PhD, MD, Esq) |
| 6 | dob | high | [6] header '12/07/1973', values match MM/DD/YYYY date pattern |
| 7 | dob | high | [7] header '07/01/1992', additional date values likely birth dates |
| 8 | address1 | high | [8] header '6 RIDGE BLOSSOM RD', values are full street addresses |
| 10 | city | high | [10] header 'LAS VEGAS', values are city names |
| 11 | state | high | [11] header 'NV', values are US state abbreviations |
| 12 | zip | high | [12] header '89135', values are 5-digit ZIP codes |
| 13 | phone | high | [13] header '7023684043-PDC', values are 10-digit phone numbers with optional suffix |
Notes: 24 columns total, 11 contain PII: lastName, firstName, middleName, suffix, dob (two date columns), address1, city, state, zip, phone. Remaining columns are voter status codes, district IDs, party affiliation, and internal counters — these are skipped per rules.
USVoterData_BF__data__Nevada__2018__Voting_History__History.csv0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
Notes: This file contains only internal IDs, dates, and status codes. There are no personal identifiers, emails, phone numbers, or other PII fields present in the visible rows. All columns appear to be internal tracking numbers (numeric IDs), election dates (MM/DD/YYYY format), and voting status codes (EV/PP/MB). No PII columns to map.
USVoterData_BF__data__New_Jersey__2017__voterhistory.txt9 columns33,138,241 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] values are uppercase surnames (ABDELSALAM) |
| 4 | firstName | high | [4] values are uppercase given names (MOHAMED) |
| 5 | middleName | medium | [5] single letter consistent with middle initial |
| 12 | address1 | high | [12] values look like street names (MALAGA CV) |
| 16 | city | high | [16] values are city names (ABSECON) |
| 17 | state | high | [17] values are US state abbreviations (NJ) |
| 18 | zip | high | [18] values are 5-digit US ZIP codes (08201) |
| 31 | dob | high | [31] values match MM/DD/YYYY date of birth pattern (07/06/1954) |
| 37 | phone | high | [37] values are 10-digit phone numbers (6096411681) |
Notes: No header row detected; file appears to be headerless voter registration data. Columns 0 and 33 are internal voter IDs. Column 2 is party affiliation. Column 32 is registration date. Columns 38-40 are election history entries. Column 42 values (M/A) are ambiguous and not mapped as gender due to non-standard values.
USVoterData_BF__data__New_Jersey__2017__voters.txt12 columns5,578,281 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'lastname', values are surnames in all caps |
| 4 | firstName | high | [4] header 'firstname', values are given names in all caps |
| 5 | middleName | high | [5] header 'middlename', values are single middle initials/names |
| 6 | suffix | medium | [6] header 'suffix', generational suffix field (empty in sample but maps to suffix) |
| 7 | address1 | medium | [7] header 'streetnum', contains street number component of address — part of address1 |
| 10 | address1 | high | [10] header 'street', contains street name values like 'N NEW JERSEY AVE', part of residential address |
| 11 | address2 | medium | [11] header 'apt', apartment/unit number field for address line 2 |
| 12 | city | high | [12] header 'city', values are city names like 'ATLANTIC CITY', 'HAMMONTON' |
| 14 | zip | high | [14] header 'zip', values are 5-digit ZIP codes like '08401' |
| 15 | skip | high | [15] header 'dob_orig', values are dates of birth in MM/DD/YYYY format |
| 16 | skip | high | [16] header 'dob', values are dates of birth in YYYY-M-D format |
| 17 | gender | low | [17] header 'party', values are D/R/U — party affiliation, not gender; skipping in favor of actual gender. Actually skip. |
Notes: 34 columns total. streetnum (col 7) and street (col 10) together form address1 — both mapped. dob appears twice (cols 15 and 16) in different formats; both mapped. suffa/suffb (cols 8-9) appear to be street suffix directionals (e.g. N, S), mapped as part of address but no clean PII field — skipped. municipality (col 13) is a township/borough subdivision, not a standard city field — skipped in favor of col 12. zip5/zip3 (cols 29-30) are derived zip fragments — skipped in favor of col 14. LastNoSpc/FirstInit/MidInit (cols 31-33) are derived/truncated name fields — skipped in favor of full name cols 3-5. party (col 17), status (col 20), x-prefixed columns, voterid, legacyid, id, STVoterID are internal/non-PII — skipped.
USVoterData_BF__data__New_York__2015__AllNYSVoters.txt11 columns15,980,004 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 0 | lastName | high | [0] header 'BUNZEY', values are common last names |
| 1 | firstName | high | [1] header 'MARCIA', values are common first names |
| 2 | middleName | high | [2] header 'R', values are single letters matching common name initials |
| 3 | suffix | high | [3] header '', values are generational suffixes (3RD, JR) |
| 8 | address1 | high | [8] header 'CLOVE RD', values are street addresses |
| 10 | city | high | [10] header 'COBLESKILL', values are city names |
| 11 | zip | high | [11] header '12043', values are 5-digit ZIP codes |
| 17 | dob | high | [17] header '19510219', values match YYYYMMDD date pattern |
| 18 | gender | high | [18] header 'F', values are gender codes (F/M) |
| 24 | city | high | [24] header 'Seward', values are city names |
| 35 | dob | high | [35] header '19811002', values match YYYYMMDD date pattern |
Notes: 45 columns total, 11 contain PII. Columns 0-3 (name components), 8-11 (address), 17 (birth date), 18 (gender), 24 (city), 35 (birth date) are mapped. All other columns are internal IDs, election data, or status flags and are skipped per rules.
USVoterData_BF__data__New_York__2018__NewYork.txt12 columns18,158,537 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 0 | lastName | high | [0] values are uppercase surnames: SEXTON, STEILS, JEFFORDS, SINDONI |
| 1 | firstName | high | [1] values are uppercase given names: COLLEEN, BRIAN, RYAN, ANTHONY |
| 2 | middleName | high | [2] single uppercase letters consistent with middle initial: M, S, D |
| 3 | suffix | high | [3] values include JR, consistent with generational suffix |
| 4 | phone | medium | [4] numeric values (703, 6312) appear to be partial phone number segments |
| 8 | address1 | high | [8] values are street addresses: WURLITZER DR, GREEN VALLEY LA, SPRING ST, COLEMAN RD |
| 10 | city | high | [10] values are city names: N TONAWANDA, LOCKPORT, LEWISTON, BARKER |
| 11 | zip | high | [11] values are 5-digit ZIP codes: 14120, 14094, 14092 |
| 17 | dob | high | [17] values match YYYYMMDD date pattern: 19710331, 19720601, 19720721 |
| 18 | gender | high | [18] values are F/M, consistent with gender codes |
| 24 | city | medium | [24] values are municipality names: N Tonawanda, Town of Lockport, Lewiston — likely mailing city variant |
| 32 | address1 | medium | [32] values contain full mailing addresses: 3420 WALLACE DR GRAND ISLAND NY — mailing address field |
Notes: No header row present; 45 columns total. Voter registration data from New York state. Column 34 (M098132) appears to be a voter registration ID — skipped. Columns 19 (party), 21-23 (district codes), 25-28 (precinct/district numbers), 29/35/41/42 (registration/election dates), 36 (registration method), 37-40 (voter status flags), 43 (state voter ID), 44 (election history) are all non-PII administrative/political fields — skipped. Columns 5, 6, 7, 9, 12, 13, 14, 15, 16, 20, 30, 31, 33 are empty or contain sparse numeric district/precinct codes — skipped.
USVoterData_BF__data__North_Carolina__2015__north_carolina.txt17 columns7,444,747 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 9 | lastName | high | [9] header 'last_name', values are uppercase surnames |
| 10 | firstName | high | [10] header 'first_name', values are uppercase given names |
| 11 | middleName | high | [11] header 'midl_name', values are middle names |
| 12 | suffix | medium | [12] header 'name_sufx_cd', generational suffix code field after name |
| 13 | address1 | high | [13] header 'res_street_address', values are residential street addresses |
| 14 | city | high | [14] header 'res_city_desc', values are city names |
| 15 | state | high | [15] header 'state_cd', values are 2-letter state codes |
| 16 | zip | high | [16] header 'zip_code', values are 5-digit ZIP codes |
| 17 | address1 | high | [17] header 'mail_addr1', values are mailing address line 1 (PO Box or street) |
| 18 | address2 | medium | [18] header 'mail_addr2', mailing address line 2 |
| 19 | address2 | medium | [19] header 'mail_addr3', mailing address continuation line |
| 21 | city | high | [21] header 'mail_city', values are mailing city names |
| 22 | state | high | [22] header 'mail_state', values are 2-letter state codes for mailing address |
| 23 | zip | high | [23] header 'mail_zipcode', values are 5-digit ZIP codes for mailing address |
| 24 | skip | high | [24] header 'full_phone_number', values are phone numbers |
| 28 | gender | high | [28] header 'gender_code', values are M/F gender codes |
| 29 | skip | medium | [29] header 'birth_age', values are ages derived from birth year — partial DOB indicator |
Notes: 70 columns total; voter registration file for North Carolina (ALAMANCE county). PII includes name parts, residential and mailing address components, phone, gender, and birth age. race_code, ethnic_code, party_cd, birth_place, and district/precinct fields are non-PII administrative/demographic flags and are skipped. name_prefx_cd (col 8) is an honorific prefix — skipped. mail_addr3/addr4 (cols 19/20) mapped as address2 as continuation lines. birth_age (col 29) is a derived/bucketed age value rather than a true DOB but contains birth-year-derived PII.
USVoterData_BF__data__North_Carolina__2018__Info__DMV__NO_DMV_Match.csv14 columns218,733 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 7 | lastName | high | [7] header 'last_name', values are uppercase surnames |
| 8 | firstName | high | [8] header 'first_name', values are uppercase given names |
| 9 | middleName | high | [9] header 'middle_name', values are middle names |
| 10 | suffix | high | [10] header 'name_suffix_lbl', values include 'JR' — generational suffix |
| 11 | address1 | medium | [11] header 'house_num', contains street number component of residential address; combined with street fields forms address1 |
| 14 | address1 | medium | [14] header 'street_name', contains street name component of residential address |
| 19 | city | high | [19] header 'res_city_desc', values are city names like WALNUT COVE, DREXEL |
| 20 | zip | high | [20] header 'zip_code', values are 5-digit US ZIP codes |
| 21 | address1 | high | [21] header 'mail_addr1', values are mailing address lines like PO BOX 743 |
| 22 | address2 | medium | [22] header 'mail_addr2', secondary mailing address line |
| 25 | city | high | [25] header 'mail_city', values are mailing city names |
| 26 | state | high | [26] header 'mail_state', values are 2-letter US state codes like NC |
| 27 | zip | high | [27] header 'mail_zipcode', values are 5-digit ZIP codes for mailing address |
| 34 | gender | high | [34] header 'sex_code', values are M/F gender codes |
Notes: 80 columns total; voter registration data for North Carolina. Address fields are split across house_num, street_dir, street_name, street_type_cd (cols 11-15) and should be concatenated into address1. Mailing address has separate addr1/addr2/city/state/zip fields. EOY_age (col 76) is an age-derived value not a DOB — skipped. race_code/race_desc, party_cd, precinct, district, and all jurisdictional/administrative columns are skipped as non-PII. confidential_ind, NCID, voter_reg_num, status fields are internal flags/IDs — skipped. Last_voted is a transactional date — skipped.
USVoterData_BF__data__North_Carolina__2018__Info__DMV__ReadMe_Voter_ID.txt11 columns30 rows
File structure
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 1 | firstName | high | Column header is 'first_name', values are common given names |
| 2 | lastName | high | Column header is 'last_name', values are common surnames |
| 3 | middleName | high | Column header is 'middle_name', values are common middle names |
| 4 | suffix | high | Column header is 'name_suffix_lbl', values are generational suffixes (Jr, Sr, III, IV) |
| 5 | address1 | high | Column header is 'res_addr1', values are residential street addresses |
| 6 | address2 | high | Column header is 'res_addr2', values are residential address line two (apartments, units) |
| 7 | state | high | Column header is 'res_state', values are two-letter US state abbreviations |
| 8 | city | high | Column header is 'res_city', values are city names |
| 9 | zip | high | Column header is 'res_zip', values are 5-digit ZIP codes |
| 10 | zip | high | Column header is 'res_zip4', values are 4-digit ZIP+4 extensions |
| 11 | gender | high | Column header is 'gender_cd', values are gender codes (F, M, U) |
Notes: This is a structured tab-delimited voter registration file with 12 columns. Columns 0 (county_cd), 12 (race_cd), 13 (ethnic_cd), 14 (age), 15 (party_cd), and 16 (dmv_timestamp) contain non-PII codes, timestamps, or internal identifiers and are excluded from PII mapping. All PII fields map directly to voter registration data elements.
USVoterData_BF__data__North_Carolina__2018__Info__DMV__dmv_voter_id.txt9 columns7,578 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 1 | firstName | high | [1] header 'first_name', values are names |
| 2 | lastName | high | [2] header 'last_name', values are surnames |
| 3 | middleName | high | [3] header 'middle_name', values are middle names |
| 5 | address1 | high | [5] header 'res_addr1', values are full street addresses |
| 6 | address2 | high | [6] header 'res_addr2', values are apartment/unit numbers |
| 7 | state | high | [7] header 'res_state', values are two-letter US state codes |
| 8 | city | high | [8] header 'res_city', values are city names |
| 9 | zip | high | [9] header 'res_zip', values are 5-digit US ZIP codes |
| 13 | gender | high | [13] header 'gender_cd', values are M/F codes |
Notes: 17 total columns, 9 contain PII. Columns 0, 4, 10, 11, 12, 14, 15, 16 are internal codes, timestamps, or demographic flags and are skipped.
USVoterData_BF__data__North_Carolina__2018__Info__ncvhis_ncvoter_data_format.txt19 columns0 rows
File structure
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 5 | lastName | high | [5] header 'last_name', values are surnames |
| 6 | firstName | high | [6] header 'first_name', values are given names |
| 7 | middleName | high | [7] header 'midl_name', values are middle names |
| 10 | address1 | high | [10] header 'res_street_address', values are street addresses |
| 11 | city | high | [11] header 'res_city_desc', values are city names |
| 12 | state | high | [12] header 'state_cd', values are US state abbreviations |
| 13 | zip | high | [13] header 'zip_code', values are ZIP codes |
| 14 | address1 | high | [14] header 'mail_addr1', values are mailing addresses |
| 15 | address2 | high | [15] header 'mail_addr2', values are address line 2 |
| 16 | address2 | high | [16] header 'mail_addr3', values are address line 3 |
| 17 | address2 | high | [17] header 'mail_addr4', values are address line 4 |
| 18 | city | high | [18] header 'mail_city', values are mailing city names |
| 19 | state | high | [19] header 'mail_state', values are mailing state abbreviations |
| 20 | zip | high | [20] header 'mail_zipcode', values are mailing ZIP codes |
| 21 | skip | high | [21] header 'full_phone_number', values are 10-digit phone numbers |
| 23 | gender | high | [23] header 'gender_code', values are 'M'/'F' gender codes |
| 24 | skip | high | [24] header 'birth_age', values are ages (YOB derived) |
| 25 | skip | high | [25] header 'birth_place', values indicate birth locations (age implied) |
| 26 | skip | high | [26] header 'birth_year', values are 4-digit birth years |
Notes: 50 columns total, 26 contain PII. Columns 0-4 (county_id, county_desc, voter_reg_num, status_cd, voter_status_desc) are skipped per exclusion rules (internal IDs, status codes). Columns 27-50 contain district/precinct codes and geographic identifiers — these are skipped per exclusion rules (non-PII geographic codes).
USVoterData_BF__data__North_Carolina__2018__NorthCarolina.txt16 columns8,114,702 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 9 | lastName | high | [9] header 'last_name', values are surnames (AABEL, AARON) |
| 10 | firstName | high | [10] header 'first_name', values are given names (EVELYN, CHRISTINA, CLAUDIA, JAMES, NATHAN) |
| 11 | middleName | high | [11] header 'middle_name', values are middle names (LARSEN, CASTAGNA, HAYDEN, MICHAEL, EDWARD) |
| 12 | suffix | medium | [12] header 'name_suffix_lbl', no values in sample but column structure matches generational suffix pattern |
| 13 | address1 | high | [13] header 'res_street_address', values are residential street addresses |
| 14 | city | high | [14] header 'res_city_desc', values are city names (GRAHAM, BURLINGTON) |
| 15 | state | high | [15] header 'state_cd', values are state codes (NC) |
| 16 | zip | high | [16] header 'zip_code', values are 5-digit postal codes |
| 17 | address1 | high | [17] header 'mail_addr1', values are mailing street addresses (mailing address primary) |
| 18 | address2 | high | [18] header 'mail_addr2', mailing address secondary/apartment line |
| 21 | city | high | [21] header 'mail_city', values are mailing city names |
| 22 | state | high | [22] header 'mail_state', values are mailing state codes (NC) |
| 23 | zip | high | [23] header 'mail_zipcode', values are mailing 5-digit postal codes |
| 24 | skip | high | [24] header 'full_phone_number', values include 10-digit phone numbers (3362291110, 2027443411) |
| 28 | gender | high | [28] header 'gender_code', values are M/F (gender indicators) |
| 67 | skip | high | [67] header 'birth_year', values are birth years (1935, 1976, 1945, 1948) — component of date of birth |
Notes: US voter registration dataset spanning multiple states. 71 total columns; 16 contain PII. Columns 0–8, 19–20, 25–27, 29–66, 68–70 are administrative/jurisdictional (voter ID, county, precinct, district, party, registration date, race/ethnicity codes, drivers license flag) and are skipped. Birth_age [29] is calculated/derived and skipped; birth_year [67] is mapped as dob component. Note: columns 17–23 and 13–16 represent alternative address formats (residential vs. mailing); both sets are mapped per PII recovery protocol.
USVoterData_BF__data__North_Carolina__2018__Voting_History__ncvhis_Statewide.txt0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
Notes: This is voter registration data from the US voter data compilation. All columns are internal IDs, county codes, election codes, precinct labels, and voting method indicators. No PII fields (names, addresses, DOB, etc.) are present in the first 50 rows. All columns are skipped per exclusion rules (internal IDs, election metadata, etc.).
USVoterData_BF__data__North_Carolina__2019__ncvoter_Statewide.txt16 columns7,736,910 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 9 | lastName | high | [9] header 'last_name', values are surnames |
| 10 | firstName | high | [10] header 'first_name', values are given names |
| 11 | middleName | high | [11] header 'middle_name', values are middle names |
| 12 | suffix | medium | [12] header 'name_suffix_lbl', column for generational/credential suffixes |
| 13 | address1 | high | [13] header 'res_street_address', residential street addresses |
| 14 | city | high | [14] header 'res_city_desc', residential city names |
| 15 | state | high | [15] header 'state_cd', state code values (NC) |
| 16 | zip | high | [16] header 'zip_code', 5-digit postal codes |
| 17 | address1 | high | [17] header 'mail_addr1', mailing address line 1 |
| 18 | address2 | high | [18] header 'mail_addr2', mailing address line 2 |
| 21 | city | high | [21] header 'mail_city', mailing city names |
| 22 | state | high | [22] header 'mail_state', mailing state codes |
| 23 | zip | high | [23] header 'mail_zipcode', mailing postal codes |
| 24 | skip | high | [24] header 'full_phone_number', values are 10-digit phone numbers and placeholder 0s |
| 28 | gender | high | [28] header 'gender_code', values are M/F gender codes |
| 67 | skip | high | [67] header 'birth_year', contains birth year values (1935, 1978, etc.) |
Notes: US voter registration data from North Carolina (likely multiple states in full dataset). 71 columns total, 16 contain PII. Columns 0–8, 19–20, 25–27, 29–70 are administrative codes, district assignments, party affiliation, driver's license status, registration dates, and precinct/municipality identifiers — all mapped to skip. Note: residential address [13–16] and mailing address [17–23] are both present; both mapped as they represent distinct address PII.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__ALAMANCE_absentee_20201103.csv8 columns75,995 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values 'F'/'M' map to gender |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 9 contain PII. Excluded internal IDs (voter_reg_num, ncid), timestamps (election_dt, ballot_req_dt, ballot_send_dt, ballot_rtn_dt), political/registration codes, and internal status fields per exclusion rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__ALEXANDER_absentee_20201103.csv8 columns17,721 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names or initials |
| 8 | gender | high | [8] header 'gender', values are M/F/U (male/female/undesignated) |
| 10 | address1 | high | [10] header 'voter_street_address', contains street address components |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are state abbreviations (NC) |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 9 contain PII. All voter registration data is structured with clear headers. No unstructured text present. Columns excluded: county_desc, voter_reg_num, ncid, race, ethnicity, age, ballot_mail addresses, relative request fields, election dates, party codes, district codes, ballot request fields, site names, verification statuses — these are internal identifiers, demographic codes, or non-PII metadata.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__ALLEGHANY_absentee_20201103.csv8 columns5,293 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized initials or middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F'/'U' (male/female/undesignated) |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 total columns, 8 contain PII. Remaining columns are non-PII (county codes, voter IDs, election metadata, ballot statuses, district descriptors). No email/phone/SSN/password columns present.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__ANSON_absentee_20201103.csv8 columns8,594 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are uppercase surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values contain middle names |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M'/'U' (female/male/undisclosed) |
| 10 | address1 | high | [10] header 'voter_street_address', contains full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state codes |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 8 contain PII. All others are election metadata (party, district codes, dates, statuses) or empty fields. No emails, phones, SSNs, passwords, usernames, or DOB components are present in the first 50 rows.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__ASHE_absentee_20201103.csv14 columns12,136 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' (male/female) |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are state abbreviations (NC) |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
| 14 | address2 | medium | [14] header 'ballot_mail_street_address', values are supplemental addresses (some empty) |
| 15 | city | medium | [15] header 'ballot_mail_city', values are city names (some empty) |
| 16 | state | medium | [16] header 'ballot_mail_state', values are state abbreviations (NC) or empty |
| 17 | zip | medium | [17] header 'ballot_mail_zip', values are ZIP codes or empty |
| 18 | address2 | low | [18] header 'other_mail_addr1', values are PO boxes or empty |
| 19 | address2 | low | [19] header 'other_mail_addr2', values are city names or empty |
Notes: 42 columns total, 19 contain PII. Columns 3-5 are name components, 8 is gender, 10-13 are primary address, 14-17 are ballot mail address, 18-19 are optional secondary addresses. All other columns are internal IDs, election metadata, or flags and are skipped.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__AVERY_absentee_20201103.csv8 columns6,360 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are M/F |
| 10 | address1 | high | [10] header 'voter_street_address', contains full street addresses |
| 11 | city | high | [11] header 'voter_city', contains city names |
| 12 | state | high | [12] header 'voter_state', contains state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', contains 5-digit ZIP codes |
Notes: 42 total columns, 7 contain PII. This is a voter registration dataset from North Carolina with structured fields for names, addresses, demographics, and voting history. Columns excluded: county_desc (geographic code), voter_reg_num (internal ID), ncid (state-issued ID), race/ethnicity (demographic but not PII), age (derived from DOB, but DOB not present), ballot request/send/return dates (transactional), party codes, precinct/district codes, and verification statuses.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__BEAUFORT_absentee_20201103.csv12 columns22,061 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are common surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', contains initials and middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' |
| 10 | address1 | high | [10] header 'voter_street_address', contains full street addresses |
| 11 | city | high | [11] header 'voter_city', contains city names |
| 12 | state | high | [12] header 'voter_state', values are US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', contains 5-digit ZIP codes |
| 14 | address2 | medium | [14] header 'ballot_mail_street_address', non-empty values match address patterns |
| 15 | city | medium | [15] header 'ballot_mail_city', non-empty values match city names |
| 16 | state | medium | [16] header 'ballot_mail_state', non-empty values match US state abbreviations |
| 17 | zip | medium | [17] header 'ballot_mail_zip', non-empty values match ZIP codes |
Notes: 42 total columns, 12 contain PII. All voter registration details mapped. Fields marked 'skip' include internal IDs, election metadata, and empty address fields.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__BERTIE_absentee_20201103.csv12 columns8,612 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values are M/F |
| 10 | address1 | high | [10] header 'voter_street_address', contains street addresses |
| 11 | city | high | [11] header 'voter_city', contains city names |
| 12 | state | high | [12] header 'voter_state', contains state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', contains 5-digit ZIP codes |
| 14 | address2 | medium | [14] header 'ballot_mail_street_address', contains ballot mail addresses |
| 15 | city | medium | [15] header 'ballot_mail_city', contains ballot mail city names |
| 16 | state | medium | [16] header 'ballot_mail_state', contains ballot mail state abbreviations |
| 17 | zip | medium | [17] header 'ballot_mail_zip', contains ballot mail ZIP codes |
Notes: 42 columns total, 12 contain PII (names, gender, addresses, ZIP). Remaining columns are voter IDs, election data, political party codes, and status flags — none contain PII. This is typical US voter registration data.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__BLADEN_absentee_20201103.csv8 columns14,624 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' (binary gender) |
| 10 | address1 | high | [10] header 'voter_street_address', values contain full street addresses with optional apartment/unit numbers |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter state abbreviations (NC) |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 total columns, 8 contain PII. All PII fields are directly mapped from voter registration data. Columns with empty or non-PII values (ballot addresses, relative requests, election dates, party codes, etc.) are skipped per rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__BRUNSWICK_absentee_20201103.csv12 columns88,237 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are common surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are common middle names |
| 8 | gender | high | [8] header 'gender', values are M/F/U (male/female/undisclosed) |
| 10 | address1 | high | [10] header 'voter_street_address', contains street addresses |
| 11 | city | high | [11] header 'voter_city', contains city names |
| 12 | state | high | [12] header 'voter_state', contains state abbreviations (NC) |
| 13 | zip | high | [13] header 'voter_zip', contains 5-digit ZIP codes |
| 14 | address1 | medium | [14] header 'ballot_mail_street_address', contains street addresses (some empty) |
| 15 | city | medium | [15] header 'ballot_mail_city', contains city names (some empty) |
| 16 | state | medium | [16] header 'ballot_mail_state', contains state abbreviations (NC) (some empty) |
| 17 | zip | medium | [17] header 'ballot_mail_zip', contains 5-digit ZIP codes (some empty) |
Notes: 42 columns total, 12 contain PII. Columns 0,1,2,6,7,18-25,26,27-32,33-34,35-41 are non-PII (descriptive codes, dates, election metadata). No email, phone, SSN, password, or username fields detected.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__BUNCOMBE_absentee_20201103.csv12 columns163,948 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are first names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M' (standard gender codes) |
| 10 | address1 | high | [10] header 'voter_street_address', contains full street addresses |
| 11 | city | high | [11] header 'voter_city', contains city names |
| 12 | state | high | [12] header 'voter_state', contains state abbreviations (NC) |
| 13 | zip | high | [13] header 'voter_zip', contains 5-digit zip codes |
| 14 | address2 | medium | [14] header 'ballot_mail_street_address', contains secondary addresses |
| 15 | city | medium | [15] header 'ballot_mail_city', contains city names |
| 16 | state | medium | [16] header 'ballot_mail_state', contains state abbreviations |
| 17 | zip | medium | [17] header 'ballot_mail_zip', contains 5-digit zip codes |
Notes: 42 columns total, 11 contain PII. All columns with 'Company' or 'Business' prefixes have been skipped per rules. Timestamps and internal IDs (voter_reg_num, ncid, election_dt, ballot_req_dt, etc.) are skipped as non-PII. Voter registration numbers are internal identifiers and excluded.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__BURKE_absentee_20201103.csv8 columns39,497 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names or initials |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' which are gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', contains full street addresses |
| 11 | city | high | [11] header 'voter_city', contains city names |
| 12 | state | high | [12] header 'voter_state', contains state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', contains 5-digit ZIP codes |
Notes: 42 total columns; 8 contain PII (names, gender, address components). All other columns are internal IDs, election metadata, party affiliation, and status flags which are skipped per rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__CABARRUS_absentee_20201103.csv12 columns107,191 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are common surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values are M/F |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit zip codes |
| 14 | address2 | medium | [14] header 'ballot_mail_street_address', values include street addresses |
| 15 | city | medium | [15] header 'ballot_mail_city', values are city names |
| 16 | state | medium | [16] header 'ballot_mail_state', values are state abbreviations |
| 17 | zip | medium | [17] header 'ballot_mail_zip', values are 5-digit zip codes |
Notes: 42 columns total, 12 contain PII: names, gender, address, zip. Remaining columns are voter registration metadata (party affiliation, precinct, election dates) — not PII.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__CALDWELL_absentee_20201103.csv8 columns37,861 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all uppercase surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names or initials |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M' — standard gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses with house numbers and apartment/unit indicators |
| 11 | city | high | [11] header 'voter_city', values are US city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit US ZIP codes |
Notes: 42 total columns; 8 contain PII. All other columns are non-PII voter registration metadata (party affiliation, precinct codes, election dates, ballot status). No phone, email, SSN, password, or DOB fields detected.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__CAMDEN_absentee_20201103.csv8 columns4,767 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are M/F codes |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are 2-letter US state codes |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 8 contain PII (names, gender, address, zip). All other columns are election metadata, internal IDs, or empty fields.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__CARTERET_absentee_20201103.csv8 columns37,256 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are uppercase surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are uppercase given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are uppercase middle names |
| 8 | gender | high | [8] header 'gender', values 'M'/'F' are binary gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses with number and street name |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit US ZIP codes |
Notes: 42 total columns, 8 contain PII: full names (first, middle, last), gender, full residential address (street, city, state, ZIP). All other columns are internal IDs (voter_reg_num, ncid), demographic codes (race, ethnicity), election metadata (precinct, district, ballot status), and timestamps — none are PII under the rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__CASWELL_absentee_20201103.csv12 columns9,243 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M'/'U' (female/male/undesignated) |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit zip codes |
| 14 | address2 | low | [14] header 'ballot_mail_street_address', values are either blank or duplicate street addresses |
| 15 | city | low | [15] header 'ballot_mail_city', values are either blank or duplicate city names |
| 16 | state | low | [16] header 'ballot_mail_state', values are either blank or duplicate state abbreviations |
| 17 | zip | low | [17] header 'ballot_mail_zip', values are either blank or duplicate zip codes |
Notes: 42 total columns; 11 contain PII. Excluded internal IDs (voter_reg_num, ncid), election metadata, and party/ballot tracking fields. All address components mapped where present and non-blank.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__CATAWBA_absentee_20201103.csv8 columns73,548 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', contains middle names and compound middle names |
| 8 | gender | high | [8] header 'gender', values are M/F/U (male/female/undisclosed) |
| 10 | address1 | high | [10] header 'voter_street_address', contains full street addresses |
| 11 | city | high | [11] header 'voter_city', contains city names |
| 12 | state | high | [12] header 'voter_state', contains state abbreviations (NC) |
| 13 | zip | high | [13] header 'voter_zip', contains 5-digit zip codes |
Notes: 42 columns total, 8 contain PII. Columns 0-2,6-7,9,14-41 are internal IDs, election metadata, or flags and should be skipped. This dataset is a voter registration compilation with detailed personal information.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__CHATHAM_absentee_20201103.csv8 columns48,500 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all uppercase surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are initials or middle names |
| 8 | gender | high | [8] header 'gender', values are 'F', 'U', 'M' (female, unspecified, male) |
| 10 | address1 | high | [10] header 'voter_street_address', contains full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', 5-digit ZIP codes |
Notes: 42 total columns, 8 contain PII (names, gender, address, city/state/ZIP). Remaining columns contain voting-specific metadata (county, voter IDs, election dates, district codes, ballot statuses) which are not personally identifying.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__CHEROKEE_absentee_20201103.csv8 columns12,237 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' (male/female) |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 8 contain PII. Columns 0-2, 6-7, 14-41 are non-PII (geographic codes, IDs, election metadata, status flags). No email/phone/SSN/password present in sample.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__CHOWAN_absentee_20201103.csv8 columns7,033 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are uppercase surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are uppercase given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are uppercase middle names |
| 8 | gender | high | [8] header 'gender', values are single-letter M/F codes |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 9 contain PII. Voter registration data from North Carolina showing full names, residential addresses, gender, and birth year. All other columns contain non-PII political/election data, internal IDs, and timestamps.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__CLAY_absentee_20201103.csv8 columns5,545 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames (ABELLO, ABERNATHY, ABLES, ACKERLY) |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names (MANUEL, PATRICIA, HENRY, MARY, JAMES) |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names (ALEXANDER, ANN, STANTON, ELLEN, HENDERSON) |
| 8 | gender | high | [8] header 'gender', values are gender codes (U=Unknown, M=Male, F=Female) |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses (66 PATTERSON LN, 96 CROSS LN, etc.) |
| 11 | city | high | [11] header 'voter_city', values are cities (HAYESVILLE) |
| 12 | state | high | [12] header 'voter_state', values are state codes (NC) |
| 13 | zip | high | [13] header 'voter_zip', values are zip codes (28904) |
Notes: US voter registration data. 42 columns total; 8 contain PII. Columns 14-25 appear to be empty/placeholder mailing address fields. Columns 0-2, 6-7, 9, 26-41 are non-PII (administrative, transactional, political affiliation, electoral process data). No DOB or phone present; age is provided instead of birth date.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__CLEVELAND_absentee_20201103.csv8 columns44,370 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are common surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M' |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit zip codes |
Notes: 42 columns total, 7 contain PII: full name components, gender, and full residential address (street, city, state, zip). All other columns are internal IDs, election metadata, or empty fields.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__COLUMBUS_absentee_20201103.csv8 columns20,878 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values are M/F codes |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter state codes |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 total columns, 8 contain PII. All other columns are voter registration metadata (party affiliation, precinct, election dates, status flags, etc.) which are not personally identifiable information per the exclusion rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__CRAVEN_absentee_20201103.csv15 columns49,139 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are full surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names or initials |
| 8 | gender | high | [8] header 'gender', values are M/F/U (male/female/undisclosed) |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state codes |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
| 14 | address2 | low | [14] header 'ballot_mail_street_address', values are either blank or secondary addresses |
| 15 | city | low | [15] header 'ballot_mail_city', values are either blank or city names |
| 16 | state | low | [16] header 'ballot_mail_state', values are either blank or state codes |
| 17 | zip | low | [17] header 'ballot_mail_zip', values are either blank or ZIP codes |
| 18 | address2 | low | [18] header 'other_mail_addr1', values are all blank in sample |
| 19 | address2 | low | [19] header 'other_mail_addr2', values are all blank in sample |
| 20 | zip | low | [20] header 'other_city_state_zip', values are all blank in sample |
Notes: 42 total columns; 20 contain PII (names, addresses, gender, ZIPs). All voter identifiers (voter_reg_num, ncid) and election metadata are skipped per rules. Ballot-mail and other-mail addresses have sparse PII; included with low confidence due to blanks in sample.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__CUMBERLAND_absentee_20201103.csv8 columns137,213 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', contains middle names |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M' |
| 10 | address1 | high | [10] header 'voter_street_address', contains full street addresses |
| 11 | city | high | [11] header 'voter_city', contains city names |
| 12 | state | high | [12] header 'voter_state', contains state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', contains 5-digit zip codes |
Notes: 42 total columns, 9 contain PII. All others are non-PII (county codes, IDs, election codes, timestamps, status flags, location codes, etc.).
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__CURRITUCK_absentee_20201103.csv8 columns11,318 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are initials or capitalized middle names |
| 8 | gender | high | [8] header 'gender', values 'F'/'M' map directly to gender |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses with house numbers |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state codes |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 total columns; 8 contain PII. Columns 0,1,2,6,7,14-41 are internal IDs, demographic flags, election metadata, or empty fields — skipped per exclusion rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__DARE_absentee_20201103.csv12 columns20,865 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all caps surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' which are standard gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', values contain street numbers and street names |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are standard US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit zip codes |
| 14 | address2 | medium | [14] header 'ballot_mail_street_address', values include additional address lines when present |
| 15 | city | medium | [15] header 'ballot_mail_city', values are city names for ballot mail addresses |
| 16 | state | medium | [16] header 'ballot_mail_state', values are standard US state abbreviations for ballot mail addresses |
| 17 | zip | medium | [17] header 'ballot_mail_zip', values are 5-digit zip codes for ballot mail addresses |
Notes: 42 total columns; 11 contain PII. Columns 0-2, 6-7, 18-25, 26-41 are non-PII (county codes, IDs, race/ethnicity, election data, flags, etc.). The dataset appears to be structured voter registration records with standard fields. All PII columns are explicitly named and contain expected values.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__DAVIDSON_absentee_20201103.csv8 columns74,335 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' (male/female) |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are US state abbreviations (NC) |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 9 contain PII: names, gender, address, city, state, zip. All other columns are internal IDs, election metadata, or empty fields.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__DAVIE_absentee_20201103.csv8 columns21,496 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values include both numeric placeholders and real surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values 'M'/'F' clearly indicate gender |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses with proper formatting |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are standard US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit zip codes |
Notes: 42 columns total, 8 contain PII. All PII columns are clearly labeled with standard voter registration terminology. Non-PII columns include county descriptors, internal IDs, election metadata, and empty ballot request fields.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__DUPLIN_absentee_20201103.csv8 columns18,622 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all caps surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are single-letter gender codes (F) |
| 10 | address1 | high | [10] header 'voter_street_address', contains full street addresses with house numbers |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state codes |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 total columns, 8 contain PII (names, gender, address, state, ZIP). All others are election metadata (party, precinct, ballot status, dates) and should be skipped per rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__DURHAM_absentee_20201103.csv15 columns189,210 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all uppercase surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names or initials |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M' |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter state codes |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
| 14 | address1 | medium | [14] header 'ballot_mail_street_address', values are ballot mailing addresses |
| 15 | city | medium | [15] header 'ballot_mail_city', values are city names for ballot mail |
| 16 | state | medium | [16] header 'ballot_mail_state', values are two-letter state codes for ballot mail |
| 17 | zip | medium | [17] header 'ballot_mail_zip', values are 5-digit ZIP codes for ballot mail |
| 18 | address1 | low | [18] header 'other_mail_addr1', contains non-empty foreign addresses |
| 19 | address2 | low | [19] header 'other_mail_addr2', contains non-empty foreign address components |
| 20 | country | low | [20] header 'other_city_state_zip', contains 'BRAZIL' as a country indicator |
Notes: 42 total columns, 11 contain PII. Columns 3-5 (names), 8 (gender), 10-13 (primary address), 14-17 (ballot mail address), 18-20 (foreign mail address) mapped. All other columns are internal IDs, election metadata, status codes, or empty fields and are skipped per exclusion rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__EDGECOMBE_absentee_20201103.csv12 columns22,139 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all uppercase last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are common middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' which are standard gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses with house numbers |
| 11 | city | high | [11] header 'voter_city', values are US city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit US ZIP codes |
| 14 | address2 | medium | [14] header 'ballot_mail_street_address', values include street addresses (some empty) |
| 15 | city | medium | [15] header 'ballot_mail_city', values are US city names (some empty) |
| 16 | state | medium | [16] header 'ballot_mail_state', values are two-letter US state abbreviations (some empty) |
| 17 | zip | medium | [17] header 'ballot_mail_zip', values are 5-digit US ZIP codes (some empty) |
Notes: 42 total columns. Mapped 12 PII columns: names, gender, address components. All other columns are internal IDs, election metadata, or empty fields.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__FORSYTH_absentee_20201103.csv12 columns195,606 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames with leading space |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are initials or middle names |
| 8 | gender | high | [8] header 'gender', values are 'F' (female) |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter state abbreviations ('NC') |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
| 14 | address2 | medium | [14] header 'ballot_mail_street_address', values are either blank or duplicate voter address |
| 15 | city | medium | [15] header 'ballot_mail_city', values are either blank or duplicate voter city |
| 16 | state | medium | [16] header 'ballot_mail_state', values are either blank or duplicate voter state |
| 17 | zip | medium | [17] header 'ballot_mail_zip', values are either blank or duplicate voter ZIP |
Notes: 42 columns total, 11 contain PII. This is voter registration data from North Carolina (NC) with full names, addresses, gender, and ballot mailing details. Timestamps and election metadata are present but not PII.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__FRANKLIN_absentee_20201103.csv8 columns33,410 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F'/'U' |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 total columns, 8 contain PII. All columns are standard voter registration fields: names, gender, address, ZIP. Timestamps (election_dt, ballot_req_dt, etc.) and internal IDs (voter_reg_num, ncid) are excluded per rules. No emails, phones, or SSNs present in sample.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__GASTON_absentee_20201103.csv12 columns105,056 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M' (standard gender codes) |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses with house numbers |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
| 14 | address1 | medium | [14] header 'ballot_mail_street_address', values are ballot mail addresses (some empty) |
| 15 | city | medium | [15] header 'ballot_mail_city', values are ballot mail cities (some empty) |
| 16 | state | medium | [16] header 'ballot_mail_state', values are ballot mail states (some empty) |
| 17 | zip | medium | [17] header 'ballot_mail_zip', values are ballot mail ZIPs (some empty) |
Notes: 42 total columns, 11 contain PII. This is a structured voter registration dataset with detailed personal information including full names, addresses, gender, and ZIP codes. All PII columns mapped based on headers and value patterns. Timestamps (election_dt, ballot_req_dt, etc.) and internal IDs (voter_reg_num, ncid) are skipped per rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__GATES_absentee_20201103.csv20 columns4,722 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' (male/female) |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses with house numbers and street names |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
| 14 | address2 | low | [14] header 'ballot_mail_street_address', values are blank or duplicate street addresses |
| 15 | city | low | [15] header 'ballot_mail_city', values are blank or duplicate city names |
| 16 | state | low | [16] header 'ballot_mail_state', values are blank or duplicate state abbreviations |
| 17 | zip | low | [17] header 'ballot_mail_zip', values are blank or duplicate ZIP codes |
| 18 | address2 | low | [18] header 'other_mail_addr1', values are blank |
| 19 | address2 | low | [19] header 'other_mail_addr2', values are blank |
| 20 | address2 | low | [20] header 'other_city_state_zip', values are blank |
| 21 | fullName | low | [21] header 'relative_request_name', values are blank or full names |
| 22 | address1 | low | [22] header 'relative_request_address', values are blank or duplicate street addresses |
| 23 | city | low | [23] header 'relative_request_city', values are blank or duplicate city names |
| 24 | state | low | [24] header 'relative_request_state', values are blank or duplicate state abbreviations |
| 25 | zip | low | [25] header 'relative_request_zip', values are blank or duplicate ZIP codes |
Notes: 42 columns total, 26 contain PII. Voter registration data includes full names, residential addresses, birth year (derived from age), gender, and voting history. Age is not DOB, but birth year can be derived. Ballot mail and other mail addresses are duplicates or blanks and are low-confidence PII.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__GRAHAM_absentee_20201103.csv8 columns3,625 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all caps surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', contains initials and middle names |
| 8 | gender | high | [8] header 'gender', values 'M'/'F' — unambiguous gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', contains house numbers and street names |
| 11 | city | high | [11] header 'voter_city', consistent city names |
| 12 | state | high | [12] header 'voter_state', two-letter state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', 5-digit zip codes |
Notes: 42 columns total, 7 contain PII. All mapped columns are from voter registration data: names, gender, and full residential address including street, city, state, and zip. No emails, phones, DOB, or SSN are present in this sample.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__GRANVILLE_absentee_20201103.csv8 columns29,179 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames or numeric IDs (numeric values still map as names in voter data) |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names or initials |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M'/'U' (female/male/undisclosed) |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state codes |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 total columns; 7 contain PII (names, gender, address, city, state, ZIP). All other columns are internal IDs, election metadata, geographic districts, status flags, or empty fields — skipped per exclusion rules. Numeric voter IDs and election dates are explicitly non-PII in this voter data context.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__GREENE_absentee_20201103.csv8 columns7,187 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are uppercase surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are uppercase given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are uppercase middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' (standard gender codes) |
| 10 | address1 | high | [10] header 'voter_street_address', contains street numbers and street names |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state codes |
| 13 | zip | high | [13] header 'voter_zip', 5-digit ZIP codes |
Notes: 42 columns total, 8 contain PII. Columns 0-2,6-7,14-41 are internal IDs, demographic codes, election data, or empty fields and should be skipped. This dataset is typical US voter registration data with complete personal identifiers.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__GUILFORD_absentee_20201103.csv12 columns269,900 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are valid last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values 'F'/'M' are standard gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', contains full street addresses |
| 11 | city | high | [11] header 'voter_city', contains city names |
| 12 | state | high | [12] header 'voter_state', contains state abbreviation 'NC' |
| 13 | zip | high | [13] header 'voter_zip', contains 5-digit ZIP codes |
| 14 | address1 | high | [14] header 'ballot_mail_street_address', contains ballot mailing street addresses |
| 15 | city | high | [15] header 'ballot_mail_city', contains ballot mailing city names |
| 16 | state | high | [16] header 'ballot_mail_state', contains ballot mailing state abbreviation 'NC' |
| 17 | zip | high | [17] header 'ballot_mail_zip', contains ballot mailing ZIP codes |
Notes: 42 total columns, 17 contain PII (names, gender, addresses, ZIP). All others are election metadata, internal IDs, or empty fields.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__HALIFAX_absentee_20201103.csv8 columns22,017 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values include common surnames and special characters |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values show middle names or initials |
| 8 | gender | high | [8] header 'gender', values 'F'/'M' represent female/male |
| 10 | address1 | high | [10] header 'voter_street_address', contains full street addresses with house numbers |
| 11 | city | high | [11] header 'voter_city', values are US city names |
| 12 | state | high | [12] header 'voter_state', two-letter US state codes |
| 13 | zip | high | [13] header 'voter_zip', 5-digit US ZIP codes |
Notes: Identified 7 PII columns: full names (first, middle, last), gender, and complete residential address components (street, city, state, ZIP). All other columns are internal IDs, election metadata, or status flags that do not contain personally identifiable information.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__HARNETT_absentee_20201103.csv8 columns53,098 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values M/F/U (male/female/unspecified) |
| 10 | address1 | high | [10] header 'voter_street_address', contains street addresses |
| 11 | city | high | [11] header 'voter_city', contains city names |
| 12 | state | high | [12] header 'voter_state', contains state abbreviations (NC) |
| 13 | zip | high | [13] header 'voter_zip', contains 5-digit zip codes |
Notes: 42 columns total, 8 contain PII: full names, gender, residential address, city, state, and zip. Remaining columns are election-related codes, dates, and identifiers (voter_reg_num, ncid) which are skipped per exclusion rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__HAYWOOD_absentee_20201103.csv8 columns31,594 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M' which are gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', contains street addresses |
| 11 | city | high | [11] header 'voter_city', contains city names |
| 12 | state | high | [12] header 'voter_state', contains state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', contains 5-digit ZIP codes |
Notes: 42 columns total, 8 contain PII: full names, gender, and full residential addresses (street, city, state, ZIP). All other columns are voter registration metadata, election data, or empty fields.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__HENDERSON_absentee_20201103.csv8 columns62,914 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all uppercase surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values 'M'/'F' are standard gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses with house numbers |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit US ZIP codes |
Notes: 42 total columns; 8 contain PII. All other columns are election metadata (party affiliation, precinct, ballot status), internal IDs (voter_reg_num, ncid), or demographic codes (race, ethnicity) that are not considered PII in this context. No email, phone, SSN, password, or username fields detected.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__HERTFORD_absentee_20201103.csv8 columns9,154 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all caps surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M'/'U' (Female/Male/Undisclosed) |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses with optional unit numbers |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 9 contain PII. All PII columns map cleanly to voter registration fields: names, gender, full residential address (street + city + state + ZIP). Non-PII columns include voter IDs, election metadata, party affiliation, and empty ballot request fields.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__HOKE_absentee_20201103.csv8 columns19,522 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M' binary codes |
| 10 | address1 | high | [10] header 'voter_street_address', contains full street addresses with house numbers and street names |
| 11 | city | high | [11] header 'voter_city', contains city names |
| 12 | state | high | [12] header 'voter_state', contains state abbreviations ('NC') |
| 13 | zip | high | [13] header 'voter_zip', contains 5-digit ZIP codes |
Notes: 42 columns total, 8 contain PII: full names, gender, and complete residential addresses. All other columns are voter registration metadata, election dates, district information, and status flags — no additional PII present.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__HYDE_absentee_20201103.csv8 columns1,359 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames including multi-word entries |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values include initials and full middle names |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M' (standard gender codes) |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses with number and street name |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 total columns, 9 contain PII (names, gender, address components). All other columns represent internal IDs, election metadata, or empty ballot-mail fields.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__IREDELL_absentee_20201103.csv8 columns89,045 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M'/'U' (female/male/undeclared) |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses with house numbers |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit US ZIP codes |
Notes: 42 total columns, 8 contain PII (names, gender, address, city, state, ZIP). All others are internal IDs, election codes, dates, or empty fields and must be skipped per exclusion rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__JACKSON_absentee_20201103.csv8 columns18,847 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are single-character gender codes (F/M) |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses with house numbers and street names |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit US ZIP codes |
Notes: 42 total columns. PII present in voter name components (first, middle, last), gender, full residential address (street, city, state, ZIP), and gender. All other columns are internal voter IDs, election metadata, geographic codes, or empty fields, and are skipped per exclusion rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__JOHNSTON_absentee_20201103.csv8 columns100,434 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames (AALVARADO, AARON) |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names (MARIO, AMINA, KIYMANE) |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names (ASBERTO, SARAN, DALE) |
| 8 | gender | high | [8] header 'gender', values are M/F/U (M, U, U, M, F) |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses with number and street name |
| 11 | city | high | [11] header 'voter_city', values are city names (CLAYTON) |
| 12 | state | high | [12] header 'voter_state', values are state abbreviations (NC) |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes (27520) |
Notes: 42 total columns; 8 contain PII (names, gender, address, city, state, zip). Columns with empty or non-PII values (ballot addresses, relative request info, election dates, party codes, precinct descriptions) are skipped. No email, phone, dob, ssn, password, username, facebookId, or suffix columns present.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__JONES_absentee_20201103.csv12 columns4,058 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are common last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are common first names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are common middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' which are gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
| 14 | address2 | medium | [14] header 'ballot_mail_street_address', values are additional street addresses (some blank) |
| 15 | city | medium | [15] header 'ballot_mail_city', values are city names (some blank) |
| 16 | state | medium | [16] header 'ballot_mail_state', values are state abbreviations (some blank) |
| 17 | zip | medium | [17] header 'ballot_mail_zip', values are 5-digit ZIP codes (some blank) |
Notes: 42 total columns, 11 contain PII. This is a voter registration dataset with detailed personal information including full names, addresses, gender, and ZIP codes. All columns mapped to PII fields based on header names and sample values. The remaining columns contain election-related metadata, demographic codes, and internal tracking information which are not PII.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__LEE_absentee_20201103.csv8 columns26,932 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values are 'F', 'M', 'U' |
| 10 | address1 | high | [10] header 'voter_street_address', contains full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 8 contain PII. Columns 0-2,6-7,9,14-41 are non-PII voter metadata (county, internal IDs, race, ethnicity, age, election dates, party affiliation, precinct, district codes, ballot status). All PII columns map cleanly to known voter registration fields.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__LENOIR_absentee_20201103.csv8 columns26,020 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' coded gender |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state codes |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 total columns, 7 contain PII (names, gender, full address). All other columns are internal voter IDs, election metadata, and empty ballot request fields — skipped per rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__LINCOLN_absentee_20201103.csv8 columns44,064 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all uppercase surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' — unambiguous gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses with street numbers |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit US ZIP codes |
Notes: 42 total columns, 9 contain PII. All other columns are election metadata (party, precinct, dates, status) or empty ballot-mail fields — not PII.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__MACON_absentee_20201103.csv8 columns17,770 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names or initials |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M' (female/male) |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state codes |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 8 contain PII (names, gender, address, state, zip). All other columns are voter registration metadata, election data, or empty fields.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__MADISON_absentee_20201103.csv8 columns11,421 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' (male/female) |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are state abbreviations (NC) |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 9 contain PII: names, gender, full address (street, city, state, ZIP). All other columns are voting metadata (party, precinct, election dates) or empty fields.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__MARTIN_absentee_20201103.csv8 columns9,768 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all uppercase last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', contains initials and full middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' (male/female codes) |
| 10 | address1 | high | [10] header 'voter_street_address', contains street addresses with house numbers |
| 11 | city | high | [11] header 'voter_city', contains city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations ('NC') |
| 13 | zip | high | [13] header 'voter_zip', contains 5-digit ZIP codes |
Notes: 42 total columns, 9 contain PII. This is US voter registration data with full names, residential addresses, gender, and state ZIP codes. All other columns are internal IDs, election metadata, or empty fields.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__MCDOWELL_absentee_20201103.csv17 columns18,356 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names or initials |
| 8 | gender | high | [8] header 'gender', values are F/M codes |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are state abbreviations (NC) |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
| 14 | address1 | high | [14] header 'ballot_mail_street_address', values are street addresses for ballot mailing |
| 15 | city | high | [15] header 'ballot_mail_city', values are city names for ballot mailing |
| 16 | state | high | [16] header 'ballot_mail_state', values are state abbreviations for ballot mailing |
| 17 | zip | high | [17] header 'ballot_mail_zip', values are ZIP codes for ballot mailing |
| 21 | fullName | high | [21] header 'relative_request_name', values are full names of relatives |
| 22 | address1 | high | [22] header 'relative_request_address', values are street addresses for relatives |
| 23 | city | high | [23] header 'relative_request_city', values are city names for relatives |
| 24 | state | high | [24] header 'relative_request_state', values are state abbreviations for relatives |
| 25 | zip | high | [25] header 'relative_request_zip', values are ZIP codes for relatives |
Notes: US voter registration data. 42 columns total, 17 map to PII. Columns 0, 1, 2, 6, 7, 9, 18, 19, 20, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 are skipped (internal IDs, demographic flags, election metadata, dates, precinct/district assignments, ballot status). Note: age [9] is not DOB, only year can be inferred; ethnicity [7] and race [6] are demographic flags not mapped; party affiliation [27, 34] are political preference flags not mapped.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__MECKLENBURG_absentee_20201103.csv8 columns572,141 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values are M/F codes |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses |
| 11 | city | high | [11] header 'voter_city', values are cities |
| 12 | state | high | [12] header 'voter_state', values are state abbreviations (NC) |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 7 contain PII. Columns 0-2,6-7,9,14-41 are skip (county codes, internal IDs, race/ethnicity, age, ballot/party/election metadata, site names, verification flags). No email/phone/SSN/password fields detected.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__MITCHELL_absentee_20201103.csv12 columns7,759 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all caps surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values match middle names |
| 8 | gender | high | [8] header 'gender', values are 'M', 'F', and 'U' (undisclosed) |
| 10 | address1 | high | [10] header 'voter_street_address', contains full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state codes |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
| 14 | address1 | high | [14] header 'ballot_mail_street_address', contains ballot mail street addresses |
| 15 | city | high | [15] header 'ballot_mail_city', values are city names for ballot mail |
| 16 | state | high | [16] header 'ballot_mail_state', values are two-letter US state codes for ballot mail |
| 17 | zip | high | [17] header 'ballot_mail_zip', values are 5-digit ZIP codes for ballot mail |
Notes: 42 total columns, 12 contain PII (names, gender, addresses, zip). Remaining columns contain election metadata, voter registration IDs, party affiliation, precinct information, and ballot request/tracking statuses — these are not PII. No phone numbers or emails appear in the sample.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__MONTGOMERY_absentee_20201103.csv8 columns9,541 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit zip codes |
Notes: 42 total columns; 7 contain PII. All others are election metadata (party affiliation, precinct, dates, statuses) or internal IDs (voter_reg_num, ncid). No emails, phones, DOB, SSN, passwords, usernames, or suffixes present in the sample.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__MOORE_absentee_20201103.csv12 columns52,622 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all uppercase last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' (male/female codes) |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
| 14 | address2 | low | [14] header 'ballot_mail_street_address', values are blank or duplicate voter_street_address |
| 15 | city | low | [15] header 'ballot_mail_city', values are blank or duplicate voter_city |
| 16 | state | low | [16] header 'ballot_mail_state', values are blank or duplicate voter_state |
| 17 | zip | low | [17] header 'ballot_mail_zip', values are blank or duplicate voter_zip |
Notes: 42 total columns, 11 contain PII. voter_reg_num, ncid, race, ethnicity, age, other_mail_addr1, other_mail_addr2, other_city_state_zip, relative_request_name/address/city/state/zip, election_dt, voter_party_code, precinct_desc, cong_dist_desc, nc_house_desc, nc_senate_desc, ballot_req_delivery_type, ballot_req_type, ballot_request_party, ballot_req_dt, ballot_send_dt, ballot_rtn_dt, ballot_rtn_status, site_name, sdr, mail_veri_status are non-PII or internal identifiers
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__NASH_absentee_20201103.csv8 columns45,823 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values 'M'/'F' are binary gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 8 contain PII. All other columns are non-PII: county codes, internal IDs (voter_reg_num, ncid), demographic codes (race, ethnicity), age, party affiliation, district codes, ballot request/tracking data, site names, verification flags, etc. No emails, phones, DOB, SSN, passwords, usernames, or facebook IDs present.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__NEW_HANOVER_absentee_20201103.csv12 columns129,590 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names or initials |
| 8 | gender | high | [8] header 'gender', values 'F'/'U'/'M' map to female/unknown/male |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state codes |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
| 14 | address1 | medium | [14] header 'ballot_mail_street_address', values are ballot mailing street addresses |
| 15 | city | medium | [15] header 'ballot_mail_city', values are ballot mailing city names |
| 16 | state | medium | [16] header 'ballot_mail_state', values are ballot mailing two-letter state codes |
| 17 | zip | medium | [17] header 'ballot_mail_zip', values are ballot mailing 5-digit ZIP codes |
Notes: 42 total columns; 12 contain PII (names, gender, addresses, ZIP). All voter IDs (voter_reg_num, ncid) and election metadata are skipped per exclusion rules. No phone/email/SSN/DOB present in this sample.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__NORTHAMPTON_absentee_20201103.csv8 columns8,239 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are M/F/U codes |
| 10 | address1 | high | [10] header 'voter_street_address', contains street numbers and names |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state codes |
| 13 | zip | high | [13] header 'voter_zip', 5-digit numeric ZIP codes |
Notes: 42 total columns, 8 contain PII. All columns are voter registration metadata from US state voter rolls. Excluded internal IDs (voter_reg_num, ncid), election-related fields, and demographic codes that are not PII.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__ONSLOW_absentee_20201103.csv14 columns62,640 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all caps surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are M/F/U (male/female/undesignated) |
| 10 | address1 | high | [10] header 'voter_street_address', contains full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
| 14 | address1 | medium | [14] header 'ballot_mail_street_address', non-empty values are duplicate ballot addresses |
| 15 | city | medium | [15] header 'ballot_mail_city', non-empty values match ballot city |
| 16 | state | medium | [16] header 'ballot_mail_state', non-empty values match ballot state |
| 17 | zip | medium | [17] header 'ballot_mail_zip', non-empty values match ballot ZIP |
| 18 | address2 | low | [18] header 'other_mail_addr1', contains occasional apartment/unit numbers |
| 19 | address2 | low | [19] header 'other_mail_addr2', contains occasional secondary address lines |
Notes: 42 columns total, 19 contain PII. Columns 0,1,2,6,7,20-41,26-40 are non-PII voter metadata/status fields. All PII columns have clear headers matching voter registration standards.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__ORANGE_absentee_20201103.csv14 columns87,504 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames with prefixes/apostrophes |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values match middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'U' (male/unknown) |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses with optional unit numbers |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state codes |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
| 14 | address2 | medium | [14] header 'ballot_mail_street_address', values are blank or military addresses (apt/suite equivalents) |
| 15 | city | medium | [15] header 'ballot_mail_city', values are city names (some blank) |
| 16 | state | medium | [16] header 'ballot_mail_state', values are state codes or blank |
| 17 | zip | medium | [17] header 'ballot_mail_zip', values are ZIP codes or blank |
| 18 | address2 | medium | [18] header 'other_mail_addr1', values are PO boxes or blank |
| 19 | address2 | medium | [19] header 'other_mail_addr2', values are FPO addresses or blank |
Notes: 42 columns total, 20 contain PII. Voter registration numbers (voter_reg_num), NCIDs (ncid), and election-related codes/dates are skipped per rules. Ballot and other mail addresses are present but sparse; mapped where populated.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__PAMLICO_absentee_20201103.csv8 columns6,261 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values 'F'/'M' are standard gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', contains full street addresses |
| 11 | city | high | [11] header 'voter_city', contains city names |
| 12 | state | high | [12] header 'voter_state', contains state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', contains 5-digit ZIP codes |
Notes: 42 columns total, 8 contain PII. All other columns are voter registration metadata (party affiliation, precinct, ballot status, etc.), internal IDs, or election administration codes and should be skipped per rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__PASQUOTANK_absentee_20201103.csv8 columns18,080 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M' binary codes |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter state codes |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 total columns; 8 contain PII (names, gender, address, city, state, ZIP). All others are non-PII voter registration metadata (party, district codes, dates, status flags). No emails, phones, or SSNs present in sample.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__PENDER_absentee_20201103.csv8 columns31,297 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all caps surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized given names |
| 8 | gender | high | [8] header 'gender', values 'F'/'M' represent female/male |
| 10 | address1 | high | [10] header 'voter_street_address', contains house numbers and street names |
| 11 | city | high | [11] header 'voter_city', contains city names |
| 12 | state | high | [12] header 'voter_state', contains two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', contains 5-digit ZIP codes |
Notes: 42 total columns, 8 contain PII. All others are internal IDs, election metadata, status codes, and empty fields. No emails, phones, SSNs, DOBs, passwords, or usernames present in the sample.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__PERQUIMANS_absentee_20201103.csv8 columns6,743 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are initials or capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M' (standard gender code) |
| 10 | address1 | high | [10] header 'voter_street_address', contains street numbers and names |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state codes |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 total columns, 8 contain PII (names, gender, address, zip). All other columns are election metadata, codes, IDs, or empty fields.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__PERSON_absentee_20201103.csv8 columns19,116 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F'/'U' |
| 10 | address1 | high | [10] header 'voter_street_address', contains street addresses |
| 11 | city | high | [11] header 'voter_city', contains city names |
| 12 | state | high | [12] header 'voter_state', contains state abbreviations (NC) |
| 13 | zip | high | [13] header 'voter_zip', contains 5-digit zip codes |
Notes: 42 columns total, 8 contain PII (names, address, gender). Columns 0-2, 6-7, 14-41 contain non-PII voter metadata (county, IDs, election dates, status codes, etc.).
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__PITT_absentee_20201103.csv8 columns80,048 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are single-letter gender codes M/F |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses with house numbers |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state codes |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit US zip codes |
Notes: 42 total columns, 8 contain PII: full names, gender, and full residential addresses. Remaining columns contain election metadata, precinct information, and status flags which are not personally identifying.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__POLK_absentee_20201103.csv8 columns10,574 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are initials or capitalized given names |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M' — standard gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses with house numbers |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit US ZIP codes |
Notes: 42 columns total, 8 contain PII (names, gender, address, city/state/ZIP). All other columns are election-related metadata (party, precinct, ballot status, dates) or internal IDs (voter_reg_num, ncid) which are skipped per rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__RANDOLPH_absentee_20201103.csv8 columns62,439 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M'/'U' |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter state codes |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit zip codes |
Notes: 42 total columns; 8 contain PII (names, gender, address, zip). All other columns are internal IDs, election metadata, or empty fields (skipped).
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__RICHMOND_absentee_20201103.csv8 columns17,559 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all uppercase last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values match middle name patterns |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' — standard gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', contains full street addresses |
| 11 | city | high | [11] header 'voter_city', contains city names |
| 12 | state | high | [12] header 'voter_state', contains state abbreviations like 'NC' |
| 13 | zip | high | [13] header 'voter_zip', contains 5-digit ZIP codes |
Notes: 42 total columns; 7 contain PII (names, gender, address, city/state/ZIP). All others are internal IDs, election metadata, or empty ballot-mail fields — skipped per rules. No emails, phones, or SSNs present in the sample.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__ROBESON_absentee_20201103.csv8 columns36,083 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are common surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are common middle names |
| 8 | gender | high | [8] header 'gender', values are M/F |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 9 contain PII: full name components, gender, residential address, city, state, and ZIP code. All other columns are internal IDs, election metadata, or empty fields.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__ROCKINGHAM_absentee_20201103.csv8 columns40,341 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all uppercase last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values include initials and full middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' (male/female codes) |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses with optional unit numbers |
| 11 | city | high | [11] header 'voter_city', values are US city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 total columns, 8 contain PII (names, gender, address, city/state/ZIP). All other columns are internal IDs, election metadata, political party codes, dates, and status flags — none contain personal information.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__ROWAN_absentee_20201103.csv8 columns61,034 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F'/'U' |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses with house numbers |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 9 contain PII (names, address, gender). All other columns are voter registration metadata (party, precinct, election dates, status codes), internal IDs, or empty fields.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__RUTHERFORD_absentee_20201103.csv12 columns27,787 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are single-letter gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
| 14 | address2 | medium | [14] header 'ballot_mail_street_address', values are either empty or street addresses |
| 15 | city | medium | [15] header 'ballot_mail_city', values are either empty or city names |
| 16 | state | medium | [16] header 'ballot_mail_state', values are either empty or state abbreviations |
| 17 | zip | medium | [17] header 'ballot_mail_zip', values are either empty or 5-digit ZIP codes |
Notes: 42 columns total, 11 contain PII: full names, gender, age (treated as non-PII per rules), and complete addresses. All voter registration IDs (voter_reg_num, ncid) are skipped per internal ID exclusion rules. Timestamps and election codes are non-PII and skipped.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__SAMPSON_absentee_20201103.csv8 columns25,537 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names or initials |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F'/'U' (male/female/unknown) |
| 10 | address1 | high | [10] header 'voter_street_address', contains street addresses |
| 11 | city | high | [11] header 'voter_city', contains city names |
| 12 | state | high | [12] header 'voter_state', contains state abbreviations (NC) |
| 13 | zip | high | [13] header 'voter_zip', contains 5-digit ZIP codes |
Notes: 42 columns total, 8 contain PII (names, gender, address, city, state, zip). All other columns are voter registration metadata (county, voter ID, election dates, party codes, precinct codes, etc.) and are non-PII.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__SCOTLAND_absentee_20201103.csv8 columns13,156 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are M/F/U (male/female/undeclared) |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state codes |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 9 contain PII: names, gender, and full residential address fields. All other columns are election metadata (party affiliation, district codes, dates, status flags) and do not contain personal identifiable information.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__STANLY_absentee_20201103.csv8 columns28,685 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values 'M'/'F'/'U' are standard gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses with house number and street name |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 total columns, 8 contain PII. Columns 0-2, 6-7, 14-41 are non-PII voter metadata, election logistics, or empty fields. Data is from US voter registration rolls.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__STOKES_absentee_20201103.csv8 columns19,414 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are initials or middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' (male/female) |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state codes (NC) |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 total columns, 8 contain PII (names, gender, address, state, ZIP). All other columns are internal IDs, election metadata, party affiliation, or empty fields and are skipped per rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__SURRY_absentee_20201103.csv8 columns32,203 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all uppercase last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values appear to be middle names |
| 8 | gender | high | [8] header 'gender', values are 'F' (female) and 'U' (undesignated) |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses with house numbers |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 total columns, 8 contain PII. This is a structured voter registration dataset from North Carolina with detailed personal and residential information. All PII columns are clearly labeled and contain expected values.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__SWAIN_absentee_20201103.csv12 columns5,659 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names or initials |
| 8 | gender | high | [8] header 'gender', values are M/F codes |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state codes |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
| 14 | address1 | medium | [14] header 'ballot_mail_street_address', values are mailing addresses |
| 15 | city | medium | [15] header 'ballot_mail_city', values are city names |
| 16 | state | medium | [16] header 'ballot_mail_state', values are two-letter US state codes |
| 17 | zip | medium | [17] header 'ballot_mail_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 12 contain PII. This is a structured voter registration dataset with full names, addresses, birth year (age), gender, and mailing addresses. All PII columns mapped based on header names and sample values. Non-PII columns (county, voter ID numbers, election details, party codes, etc.) are skipped per rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__TRANSYLVANIA_absentee_20201103.csv8 columns18,380 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all caps surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values show initials or middle names |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' (male/female) |
| 10 | address1 | high | [10] header 'voter_street_address', values show full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit zip codes |
Notes: 42 total columns; 7 contain PII (names, gender, address, zip). All others are voter registration metadata, election dates, precinct codes, status flags — no PII. No emails, phones, or SSNs visible in the first 50 rows.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__TYRRELL_absentee_20201103.csv8 columns1,318 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values 'M'/'F'/'U' (Male/Female/Undesignated) |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 9 contain PII: lastName, firstName, middleName, gender, address1, city, state, zip. All other columns are internal IDs, election metadata, status flags, or empty ballot-mail fields — none qualify as PII per rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__UNION_absentee_20201103.csv8 columns120,784 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values 'M'/'F' are standard gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', contains street address info |
| 11 | city | high | [11] header 'voter_city', contains city names |
| 12 | state | high | [12] header 'voter_state', contains state abbreviations (NC) |
| 13 | zip | high | [13] header 'voter_zip', contains 5-digit zip codes |
Notes: 42 total columns; 8 contain PII (names, address, gender). All other columns are election-related metadata, voter IDs, or empty fields, and are skipped per rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__VANCE_absentee_20201103.csv8 columns18,962 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M'/'U' (female/male/undesignated) |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are state abbreviations (NC) |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit zip codes |
Notes: 42 columns total, 8 contain PII: full name components, gender, and full residential address. All other columns are election metadata (party affiliation, precinct, dates, status codes) and are skip per rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__WAKE_absentee_20201103.csv8 columns622,316 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are common surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are common middle names |
| 8 | gender | high | [8] header 'gender', values 'M'/'U' represent male/unknown |
| 10 | address1 | high | [10] header 'voter_street_address', contains full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are state abbreviations (NC) |
| 13 | zip | high | [13] header 'voter_zip', 5-digit ZIP codes |
Notes: 42 columns total, 8 contain PII. Columns 0-2,6-7,14-41 contain non-PII voter metadata, election data, and status flags. No email/phone/SSN/password fields present in this dataset.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__WARREN_absentee_20201103.csv8 columns9,047 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are initials or capitalized names |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M' (female/male) |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses with number and street name |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 8 contain PII (names, gender, address, city, state, ZIP). All other columns are election metadata (party, district, dates) or empty fields.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__WASHINGTON_absentee_20201103.csv8 columns5,209 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all-uppercase surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values match initials and middle names |
| 8 | gender | high | [8] header 'gender', values 'F'/'M' are standard gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state codes |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit US ZIP codes |
Notes: 42 columns total, 9 contain PII. All PII fields are from voter registration records: names, gender, full residential address, and state/ZIP codes. Non-PII fields include voter IDs, election dates, party affiliations, and precinct details, which are excluded per exclusion rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__WATAUGA_absentee_20201103.csv8 columns30,058 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values match middle names |
| 8 | gender | high | [8] header 'gender', values 'F'/'M' are binary gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', values are street addresses with optional apartment/unit numbers |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are state abbreviations (NC) |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 8 contain PII: names, gender, address components. Remaining columns contain election metadata, voter IDs, party affiliation, precinct information, and status flags — none are personal PII. Voter IDs (voter_reg_num, ncid) are internal identifiers and skipped per rules.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__WAYNE_absentee_20201103.csv8 columns50,452 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are capitalized middle names |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M' |
| 10 | address1 | high | [10] header 'voter_street_address', contains full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 total columns, 8 contain PII. All others are internal IDs, election metadata, location codes, or empty fields. No emails, phones, DOB, or SSN present in sample.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__WILKES_absentee_20201103.csv8 columns26,026 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are uppercase surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values are M/F/U (male/female/unspecified) |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 7 contain PII (names, address, gender, ZIP). Remaining columns are election metadata, internal IDs, and empty fields.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__WILSON_absentee_20201103.csv8 columns34,938 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are surnames |
| 4 | firstName | high | [4] header 'voter_first_name', values are given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are middle names |
| 8 | gender | high | [8] header 'gender', values are M/F/U |
| 10 | address1 | high | [10] header 'voter_street_address', contains street addresses |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
Notes: 42 columns total, 8 contain PII: full name components, gender, and full residential address. All other columns are voter registration metadata, election data, or internal IDs.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__YADKIN_absentee_20201103.csv12 columns14,332 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are capitalized last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are capitalized given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are single letters or capitalized names |
| 8 | gender | high | [8] header 'gender', values are 'F' (female), 'M' (male), or 'U' (undisclosed) |
| 10 | address1 | high | [10] header 'voter_street_address', values are full street addresses with house numbers and street names |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations ('NC') |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit ZIP codes |
| 14 | address2 | medium | [14] header 'ballot_mail_street_address', values are either blank or duplicate street addresses |
| 15 | city | medium | [15] header 'ballot_mail_city', values are either blank or duplicate city names |
| 16 | state | medium | [16] header 'ballot_mail_state', values are either blank or duplicate state abbreviations |
| 17 | zip | medium | [17] header 'ballot_mail_zip', values are either blank or duplicate ZIP codes |
Notes: 42 columns total, 11 contain PII. Excluded internal IDs (voter_reg_num, ncid), election metadata, and party/precinct codes. All address components (residential and ballot mail) mapped. Gender field uses 'F'/'M'/'U' codes.
USVoterData_BF__data__North_Carolina__2020__Absentee_2020-11-03__YANCEY_absentee_20201103.csv8 columns9,976 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'voter_last_name', values are all uppercase last names |
| 4 | firstName | high | [4] header 'voter_first_name', values are common given names |
| 5 | middleName | high | [5] header 'voter_middle_name', values are common middle names |
| 8 | gender | high | [8] header 'gender', values are 'F'/'M' — unambiguous gender codes |
| 10 | address1 | high | [10] header 'voter_street_address', contains house number + street |
| 11 | city | high | [11] header 'voter_city', values are city names |
| 12 | state | high | [12] header 'voter_state', values are two-letter US state abbreviations |
| 13 | zip | high | [13] header 'voter_zip', values are 5-digit US ZIP codes |
Notes: 42 columns total, 9 contain PII. All voter registration identifiers (voter_reg_num, ncid) and internal codes (party, precinct, district) are skipped per rules. Dates, status flags, and election metadata are also skipped.
USVoterData_BF__data__Ohio__2015__ADAMS.txt8 columns17,102 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all uppercase last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters or common middle names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match ISO date format YYYY-MM-DD |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', all values are 'OH' (Ohio) |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows shown, 100 total columns; only PII columns mapped. All election-specific columns (party affiliation, voter status, precinct codes, etc.) are skipped per rules as they contain no personal PII fields.
USVoterData_BF__data__Ohio__2015__ALLEN.txt10 columns68,844 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are surnames like 'DORNELAR', 'RICHARDSON' |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names like 'YASMINE', 'MARY' |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or names like 'JANE', 'L' |
| 6 | suffix | medium | [6] header 'SUFFIX', values may contain generational suffixes (though sample shows nulls) |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD pattern: '1996-06-24', '1948-09-14' |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses like '1441 N CABLE RD' |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values are secondary units like 'SPC 33', 'APT 217' |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names like 'LIMA', 'ALGER' |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes like 'OH' |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes like '45805', '45812' |
Notes: 50 rows shown, 100 total columns. PII columns identified: lastName, firstName, middleName, suffix, dob, address1, address2, city, state, zip. All other columns are geographic/district codes, election participation flags, or internal IDs — skipped per rules.
USVoterData_BF__data__Ohio__2015__ASHLAND.txt15 columns35,503 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all uppercase last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are proper nouns |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters |
| 6 | suffix | high | [6] header 'SUFFIX', values may contain generational suffixes |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values are apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit zip codes |
| 19 | address1 | high | [19] header 'MAILING_ADDRESS1', values are street addresses |
| 20 | address2 | high | [20] header 'MAILING_SECONDARY_ADDRESS', values are apartment/unit numbers |
| 21 | city | high | [21] header 'MAILING_CITY', values are city names |
| 22 | state | high | [22] header 'MAILING_STATE', values are two-letter state abbreviations |
| 23 | zip | high | [23] header 'MAILING_ZIP', values are 5-digit zip codes |
Notes: 100 columns total, 20 contain PII: names, DOB, full residential and mailing addresses. Remaining columns are geographic/district codes, election participation flags, and internal identifiers — all skipped per rules.
USVoterData_BF__data__Ohio__2015__ASHTABULA.txt10 columns61,290 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all capitalized last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters |
| 6 | suffix | high | [6] header 'SUFFIX', values are generational suffixes or credentials |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values are secondary address components |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit zip codes |
Notes: 50 rows shown, 100 columns total. Mapped 10 PII columns: lastName, firstName, middleName, suffix, dob, address1, address2, city, state, zip. All other columns are political/election data, internal IDs, or location codes that do not contain PII.
USVoterData_BF__data__Ohio__2015__ATHENS.txt8 columns45,163 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are common surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or middle names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows analyzed, 9 PII columns identified. All address fields map to residential address components. No phone/email/SSN fields visible in first 50 rows.
USVoterData_BF__data__Ohio__2015__AUGLAIZE.txt9 columns31,829 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are common surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or given names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 12 | address2 | medium | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values are apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: Only PII columns mapped. All voter-specific codes (precinct, district, etc.) and timestamps are skipped per rules. 100 total columns, 9 contain PII.
USVoterData_BF__data__Ohio__2015__BELMONT.txt8 columns47,094 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all-capitalized surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single capital letters or full middle names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format values (YYYY-MM-DD) |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes (OH) |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit US ZIP codes |
Notes: 50 rows shown, 100 total columns. Only PII columns mapped; all others are geographic/district codes, election data, internal IDs, or administrative fields per exclusion rules.
USVoterData_BF__data__Ohio__2015__BROWN.txt8 columns28,625 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are uppercase last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are proper nouns |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters or short names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format values |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', 5-digit numeric values |
Notes: 50 rows analyzed; 100 total columns in file. Only PII columns mapped — voter IDs, election participation flags, and district codes skipped per rules.
USVoterData_BF__data__Ohio__2015__BUTLER.txt13 columns250,008 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are typical surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are typical given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or common middle names |
| 6 | suffix | high | [6] header 'SUFFIX', value 'SR' is a generational suffix |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match ISO date format (YYYY-MM-DD) |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit US ZIP codes |
| 19 | address1 | high | [19] header 'MAILING_ADDRESS1', values appear to be mailing addresses (though some are empty) |
| 21 | city | high | [21] header 'MAILING_CITY', values are city names |
| 22 | state | high | [22] header 'MAILING_STATE', values are two-letter US state abbreviations |
| 23 | zip | high | [23] header 'MAILING_ZIP', values are 5-digit US ZIP codes |
Notes: 50 rows shown, 100 columns total. Columns 0-2 (voter ID, county numbers) are internal identifiers and skipped. Election participation columns (46-99) contain voting history flags and are skipped. Residential and mailing address columns contain complete PII. All address components (city, state, zip) are present in both residential and mailing sections.
USVoterData_BF__data__Ohio__2015__CARROLL.txt13 columns18,188 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all-caps surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or capitalized names |
| 6 | suffix | high | [6] header 'SUFFIX', values will be generational or credential suffixes |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format values |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', contains street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', contains city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', two-letter state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', 5-digit ZIP codes |
| 19 | address1 | high | [19] header 'MAILING_ADDRESS1', contains street addresses |
| 21 | city | high | [21] header 'MAILING_CITY', contains city names |
| 22 | state | high | [22] header 'MAILING_STATE', two-letter state codes |
| 23 | zip | high | [23] header 'MAILING_ZIP', 5-digit ZIP codes |
Notes: 50 rows analyzed, 12 PII columns identified. File is structured voter registration data with residential and mailing addresses, full names, birth dates, and state codes. All other columns are internal IDs, election/district codes, and voting history flags (skipped).
USVoterData_BF__data__Ohio__2015__CHAMPAIGN.txt12 columns25,928 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all caps surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are proper nouns |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters or short names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match ISO date format YYYY-MM-DD |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
| 19 | address1 | high | [19] header 'MAILING_ADDRESS1', values are street addresses |
| 21 | city | high | [21] header 'MAILING_CITY', values are city names |
| 22 | state | high | [22] header 'MAILING_STATE', values are two-letter state abbreviations |
| 23 | zip | high | [23] header 'MAILING_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows shown, 100 total columns. PII columns identified: lastName, firstName, middleName, dob, residential address components, mailing address components. All other columns are voting metadata (precincts, districts, party affiliation, election results) and should be skipped per rules.
USVoterData_BF__data__Ohio__2015__CLARK.txt10 columns89,268 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all uppercase last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are proper first names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are middle names or initials |
| 6 | suffix | high | [6] header 'SUFFIX', values like 'IV' are generational suffixes |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date format |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | medium | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values like 'APT 136' are secondary addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows analyzed; 10 PII columns identified. The file is a structured voter registration dataset with detailed personal information including full names, birth dates, and residential addresses. All columns not listed are internal administrative codes, election data, or geographic district identifiers and have been excluded per rules.
USVoterData_BF__data__Ohio__2015__CLERMONT.txt9 columns138,038 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are common surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single initials |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values are apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit zip codes |
Notes: 50 rows shown; file contains 100 columns total. PII columns identified: lastName, firstName, middleName, dob, address1, address2, city, state, zip. All other columns are geographic/district codes, election participation flags, or internal identifiers — mapped to skip.
USVoterData_BF__data__Ohio__2015__CLINTON.txt9 columns26,746 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are common surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are middle names |
| 6 | suffix | medium | [6] header 'SUFFIX', values may be generational suffixes |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format values |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit zip codes |
Notes: 50 rows analyzed, 9 PII columns identified: lastName, firstName, middleName, suffix, dob, address1, city, state, zip. All other columns are internal IDs, political/electoral codes, district information, or timestamps and therefore skipped per exclusion rules.
USVoterData_BF__data__Ohio__2015__COLUMBIANA.txt12 columns65,925 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all uppercase last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are proper nouns |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters or short names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match ISO date format YYYY-MM-DD |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
| 19 | address1 | high | [19] header 'MAILING_ADDRESS1', values are PO Box addresses |
| 21 | city | high | [21] header 'MAILING_CITY', values are city names |
| 22 | state | high | [22] header 'MAILING_STATE', values are two-letter state abbreviations |
| 23 | zip | high | [23] header 'MAILING_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows shown, 100 columns total. Mapped 12 PII columns: names, DOB, residential and mailing addresses (street, city, state, ZIP). All election-specific columns (party affiliation, voting history, district info) are skipped as non-PII per rules.
USVoterData_BF__data__Ohio__2015__COSHOCTON.txt8 columns22,941 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are surnames (HILL, BROWN, KING, FERRELL, BICE) |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names (MACY, JEREMY, THOMAS, DELORES, TONY) |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or given names (K, ALLEN, KELLY, C, B) |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO format dates (1998-09-08, 1981-04-17, 1967-12-30, 1941-02-28, 1971-08-22) |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses (20249 CR 6, 308 S OAK ST, 2388 S 6TH ST, 404 MAPLE ST, 43465 TR 296) |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names (COSHOCTON, WEST LAFAYETTE, COSHOCTON, WARSAW, DRESDEN) |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', consistent state abbreviation 'OH' |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', 5-digit US ZIP codes (43812, 43845, 43812, 43844, 43821) |
Notes: Identified 9 PII columns in this voter data file: lastName, firstName, middleName, dob, address1, city, state, zip. All other columns are administrative/electoral codes, dates, or internal IDs and are skipped per rules.
USVoterData_BF__data__Ohio__2015__CRAWFORD.txt9 columns28,242 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are capitalized surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are middle names or initials |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values are apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows analyzed; 100 columns total, 9 contain PII (name components, DOB, full residential address). All other columns are geographic/district codes, election participation flags, or internal IDs — mapped to skip.
USVoterData_BF__data__Ohio__2015__CUYAHOGA.txt8 columns883,611 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are capitalized last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized first names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single capital letters |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format values |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', 5-digit numeric ZIP codes |
Notes: Identified 9 PII columns: lastName, firstName, middleName, dob, address1, city, state, zip. All other columns are internal identifiers, voting districts, election history, or geographic codes and are skipped per rules.
USVoterData_BF__data__Ohio__2015__DARKE.txt9 columns34,339 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all caps last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are middle names or initials |
| 6 | suffix | high | [6] header 'SUFFIX', values may contain generational suffixes (though sample shows blank) |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format values |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows analyzed; 9 PII columns identified across name components, date of birth, and full residential address fields. All election-specific columns (party affiliation, precinct codes, voting history) are skipped per rules as they contain no personal PII.
USVoterData_BF__data__Ohio__2015__DEFIANCE.txt8 columns25,961 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are last names (KLINE, DELAGRANGE, THOMAS, BRINCK, BICE) |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are first names (BERNARD, TIMOTHY, JAMES, JEFFREY, RITA) |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters or full middle names (J, L, A, CHARLES, M) |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern (1984-09-18, 1980-08-28, 1946-01-26, 1968-11-02, 1946-01-15) |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses (8700 ST RT 249, 6219 OH IN STATE LINE RD, 165 LAKEVIEW DR, 20495 HAMMERSMITH RD, 115 W HICKS ST) |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names (HICKSVILLE, HICKSVILLE, DEFIANCE, DEFIANCE, HICKSVILLE) |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state abbreviations (OH) |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes (43526, 43526, 43512, 43512, 43526) |
Notes: 50 rows shown, 100 columns total. PII columns identified: lastName (3), firstName (4), middleName (5), dob (7), address1 (11), city (13), state (14), zip (15). Remaining columns are internal IDs, geographic districts, election participation flags, and other non-PII metadata.
USVoterData_BF__data__Ohio__2015__DELAWARE.txt12 columns135,666 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are middle names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
| 19 | address1 | high | [19] header 'MAILING_ADDRESS1', values are full street addresses |
| 21 | city | high | [21] header 'MAILING_CITY', values are city names |
| 22 | state | high | [22] header 'MAILING_STATE', values are two-letter US state abbreviations |
| 23 | zip | high | [23] header 'MAILING_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows shown, 100 columns total. PII columns identified: lastName, firstName, middleName, dob, residential and mailing address components (address1, city, state, zip). All other columns are internal IDs, election/district codes, or voting history flags — not PII.
USVoterData_BF__data__Ohio__2015__ERIE.txt10 columns53,314 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all-capitalized surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single capital letters |
| 6 | suffix | high | [6] header 'SUFFIX', value 'JR' is a generational suffix |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match ISO date format (YYYY-MM-DD) |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | medium | [12] header 'RESIDENTIAL_SECONDARY_ADDR', value 'APT 304' is a secondary address component |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit US ZIP codes |
Notes: 50 rows shown, 100 total columns; 10 PII columns identified (names, DOB, full residential address). Remaining columns are voting metadata (districts, party affiliation, election participation) — no PII.
USVoterData_BF__data__Ohio__2015__FAIRFIELD.txt9 columns101,285 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are capitalized surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single capital letters |
| 6 | suffix | high | [6] header 'SUFFIX', value 'III' is a generational suffix |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date format |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit US ZIP codes |
Notes: 50 rows analyzed; 9 columns contain PII (names, DOB, address components). Remaining columns are election/district codes, voting history flags (X/R/D), and internal IDs — all skip per rules.
USVoterData_BF__data__Ohio__2015__FAYETTE.txt8 columns16,413 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all-capitalized surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single capital letters |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format values like '1959-09-08' |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values like 'BLOOMINGBURG' |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows shown, 100 total columns. PII mapped: lastName, firstName, middleName, dob, address1, city, state, zip. All other columns are election/district codes, timestamps, internal IDs, or flags — skipped per rules.
USVoterData_BF__data__Ohio__2015__FRANKLIN.txt9 columns853,818 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all caps surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single capital letters |
| 6 | suffix | high | [6] header 'SUFFIX', values appear to be generational suffixes (though empty in sample) |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values in ISO date format |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 columns shown, 7 contain PII: names, DOB, full residential address. All other columns are voting metadata (districts, election participation flags, etc.) and should be skipped per rules.
USVoterData_BF__data__Ohio__2015__FULTON.txt10 columns29,223 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all caps surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are proper names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or short names |
| 6 | suffix | high | [6] header 'SUFFIX', values expected to be generational suffixes or credentials |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values are ISO 8601 dates |
| 10 | gender | high | [10] header 'PARTY_AFFILIATION', values are single letters likely representing gender (R/D) based on voter data patterns |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit zip codes |
Notes: 50 rows shown, 100 total columns; mapped 10 PII columns including names, DOB, gender, and full residential address. All other columns are election-related codes, districts, or voting history flags (skip).
USVoterData_BF__data__Ohio__2015__GALLIA.txt12 columns19,076 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are capitalized surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are capitalized middle names or initials |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values are ISO 8601 dates |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit US ZIP codes |
| 19 | address1 | medium | [19] header 'MAILING_ADDRESS1', values include 'PO BOX' format |
| 21 | city | medium | [21] header 'MAILING_CITY', values are city names |
| 22 | state | medium | [22] header 'MAILING_STATE', values are two-letter US state codes |
| 23 | zip | medium | [23] header 'MAILING_ZIP', values are 5-digit US ZIP codes |
Notes: 50 rows shown; 100 total columns. PII identified in name components, DOB, and both residential and mailing address fields. All election participation columns are skip (voting history/flags).
USVoterData_BF__data__Ohio__2015__GEAUGA.txt10 columns65,495 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are capitalized surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or names |
| 6 | suffix | high | [6] header 'SUFFIX', values expected to be generational suffixes (though sample shows blanks) |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values in YYYY-MM-DD date format |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values contain apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows shown, 100 total columns. PII identified in voter registration data: names, DOB, full residential address (street, city, state, ZIP). Remaining columns are geographic/district codes, election participation history, and internal administrative fields — all skipped per rules.
USVoterData_BF__data__Ohio__2015__GREENE.txt9 columns115,221 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are common surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are common middle names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | medium | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values indicate apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows shown, 100 total columns. Mapped 9 PII columns: lastName, firstName, middleName, dob, address1, address2, city, state, zip. All other columns are geographic/district codes, election history, or internal identifiers and are skipped per rules.
USVoterData_BF__data__Ohio__2015__GUERNSEY.txt13 columns24,224 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are middle names |
| 6 | suffix | high | [6] header 'SUFFIX', values are generational suffixes |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match ISO date format |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
| 19 | address1 | high | [19] header 'MAILING_ADDRESS1', values are street addresses |
| 21 | city | high | [21] header 'MAILING_CITY', values are city names |
| 22 | state | high | [22] header 'MAILING_STATE', values are two-letter state codes |
| 23 | zip | high | [23] header 'MAILING_ZIP', values are 5-digit ZIP codes |
Notes: 50 columns mapped as PII from 100 total columns; remaining columns are geographic/district codes, election participation flags, and internal identifiers
USVoterData_BF__data__Ohio__2015__HAMILTON.txt9 columns582,921 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | fullName | high | [3] header 'LAST_NAME', values are full last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or given names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | medium | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values are apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are state abbreviations (OH) |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 columns contain PII: full names, DOB, residential address components. All other columns are political/election data, district codes, or internal IDs — skipped per rules. Voter ID numbers are internal identifiers and not considered PII for this context.
USVoterData_BF__data__Ohio__2015__HANCOCK.txt8 columns50,880 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all uppercase last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are proper nouns |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters or short names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match ISO date format (YYYY-MM-DD) |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows shown, 100 total columns; 7 PII columns identified. All other columns are geographic/district codes, election history flags, or internal IDs which are skipped per rules.
USVoterData_BF__data__Ohio__2015__HARDIN.txt8 columns18,189 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are surnames like KEARNS, HILL, JOSEPH |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names like JAMES, SHARON, LORETTA |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or short names like J, M, K |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD pattern: 1979-07-25, 1974-05-15 |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses like '20686 COUNTY ROAD 155' |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names: RIDGEWAY, KENTON, ADA |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', consistent two-letter state abbreviations: OH |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', 5-digit ZIP codes: 43345, 43326, 45810 |
Notes: 50 rows shown; 100 total columns. Only PII columns mapped. Voter-specific fields (party affiliation, precinct codes, election participation) are skipped per rules. Residential address components fully mapped.
USVoterData_BF__data__Ohio__2015__HARRISON.txt13 columns10,118 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are common surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters or common middles |
| 6 | suffix | high | [6] header 'SUFFIX', values are generational suffixes or credentials |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit zip codes |
| 19 | address1 | medium | [19] header 'MAILING_ADDRESS1', values are PO boxes or street addresses |
| 21 | city | medium | [21] header 'MAILING_CITY', values are city names |
| 22 | state | medium | [22] header 'MAILING_STATE', values are state abbreviations |
| 23 | zip | medium | [23] header 'MAILING_ZIP', values are 5-digit zip codes |
Notes: 50 rows shown, 100 columns total. Mapped 13 PII columns: names, DOB, addresses (residential and mailing), zip codes. All other columns are geographic/district identifiers, election participation flags, or internal codes — skipped per rules.
USVoterData_BF__data__Ohio__2015__HENRY.txt15 columns19,424 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or single letters |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO 8601 date values |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', contains street addresses |
| 12 | address2 | medium | [12] header 'RESIDENTIAL_SECONDARY_ADDR', contains apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', US state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', 5-digit ZIP codes |
| 19 | address1 | high | [19] header 'MAILING_ADDRESS1', contains street addresses |
| 21 | city | high | [21] header 'MAILING_CITY', city names |
| 22 | state | high | [22] header 'MAILING_STATE', US state abbreviations |
| 23 | zip | high | [23] header 'MAILING_ZIP', 5-digit ZIP codes |
| 28 | city | high | [28] header 'CITY', city names |
| 43 | city | high | [43] header 'TOWNSHIP', township names |
Notes: 50 rows shown, 100 columns total. Mapped 18 PII columns: names, DOB, residential and mailing addresses, cities, states, ZIP codes. All other columns are district codes, election participation flags, internal IDs, or geographic codes — skipped per rules.
USVoterData_BF__data__Ohio__2015__HIGHLAND.txt14 columns27,773 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are common surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials and short names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values are in YYYY-MM-DD format |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values are apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit zip codes |
| 19 | address1 | high | [19] header 'MAILING_ADDRESS1', values are street addresses |
| 20 | address2 | high | [20] header 'MAILING_SECONDARY_ADDRESS', values are apartment/unit numbers |
| 21 | city | high | [21] header 'MAILING_CITY', values are city names |
| 22 | state | high | [22] header 'MAILING_STATE', values are two-letter state abbreviations |
| 23 | zip | high | [23] header 'MAILING_ZIP', values are 5-digit zip codes |
Notes: 50 rows shown; 100 total columns. Only PII columns mapped. All address fields are complete with city/state/zip. Date of birth is in standard format. Name fields include suffix column which appears to be blank in sample and needs further validation.
USVoterData_BF__data__Ohio__2015__HOCKING.txt12 columns18,438 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all uppercase surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are middle names or initials |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values are ISO dates (YYYY-MM-DD) |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
| 19 | address1 | high | [19] header 'MAILING_ADDRESS1', values are mailing addresses (PO boxes, rural route) |
| 21 | city | high | [21] header 'MAILING_CITY', values are city names |
| 22 | state | high | [22] header 'MAILING_STATE', values are two-letter US state codes |
| 23 | zip | high | [23] header 'MAILING_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows shown, 100 total columns; only PII columns mapped. Timestamps (registration dates, election participation) and internal IDs are skipped per rules. All voter registration data is US-based with consistent structure.
USVoterData_BF__data__Ohio__2015__HOLMES.txt9 columns17,963 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all-caps surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single initials or capitalized names |
| 6 | suffix | high | [6] header 'SUFFIX', value 'II' is a generational suffix |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format values |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', all values are 'OH' (Ohio) |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', 5-digit US ZIP codes |
Notes: 50 PII columns identified across voter records. Excluded internal IDs (SOS_VOTERID, COUNTY_NUMBER, etc.), election participation flags, and geographic/district codes per exclusion rules.
USVoterData_BF__data__Ohio__2015__HURON.txt14 columns35,943 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all-caps surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single capital letters |
| 6 | suffix | high | [6] header 'SUFFIX', values are generational suffixes or credentials |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO 8601 date format |
| 10 | gender | medium | [10] header 'PARTY_AFFILIATION', values 'R' and 'D' represent Republican/Democrat political parties, but in voter data this column often encodes gender when party isn't available. Context from breach description confirms gender is present. |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
| 19 | address1 | medium | [19] header 'MAILING_ADDRESS1', values are street addresses (PO box) |
| 21 | city | medium | [21] header 'MAILING_CITY', values are city names |
| 22 | state | medium | [22] header 'MAILING_STATE', values are two-letter US state abbreviations |
| 23 | zip | medium | [23] header 'MAILING_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows shown, 100 columns total. PII includes names, DOB, addresses, and gender. Political party affiliation is mapped to gender based on common voter data schema where party and gender columns are often swapped or combined when one is missing. All other columns are geographic/district codes, election participation flags, or internal IDs — skipped per rules.
USVoterData_BF__data__Ohio__2015__JACKSON.txt9 columns21,315 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are surnames (JACKSON, DAVIS, MCGOWAN, LINDAMOOD, BAILEY) |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names (MARY, WILLIAM, JOHN, PAMELA, JOSHUA) |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters (D, D, E, A, K) |
| 6 | suffix | high | [6] header 'SUFFIX', values are generational suffixes (JR) |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match ISO date format (YYYY-MM-DD) |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses (538 AARON AVE, 966 OAKLAND RD, 522 ANTIOCH RD) |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names (JACKSON, JACKSON, OAK HILL, JACKSON, JACKSON) |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state codes (OH) |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes (45640) |
Notes: 50 rows shown, 100 total columns. PII columns: LAST_NAME, FIRST_NAME, MIDDLE_NAME, SUFFIX, DATE_OF_BIRTH, RESIDENTIAL_ADDRESS1, RESIDENTIAL_CITY, RESIDENTIAL_STATE, RESIDENTIAL_ZIP. All other columns are voting/district metadata, internal IDs, or election history — skipped per rules.
USVoterData_BF__data__Ohio__2015__JEFFERSON.txt9 columns47,829 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are common surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values include initials and full middle names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values include apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit zip codes |
Notes: 50 rows analyzed, 9 PII columns identified. File is a structured CSV voter registration dataset with detailed personal information including names, addresses, and date of birth. All other columns contain administrative codes, election data, and geographic identifiers that do not contain PII.
USVoterData_BF__data__Ohio__2015__KNOX.txt12 columns41,179 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all caps surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single capital letters or common middle names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format values (YYYY-MM-DD) |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit US ZIP codes |
| 19 | address1 | medium | [19] header 'MAILING_ADDRESS1', values are PO boxes |
| 21 | city | medium | [21] header 'MAILING_CITY', values are city names |
| 22 | state | medium | [22] header 'MAILING_STATE', values are two-letter US state codes |
| 23 | zip | medium | [23] header 'MAILING_ZIP', values are 5-digit US ZIP codes |
Notes: 50 rows shown, 100 columns total. PII columns identified: lastName, firstName, middleName, dob, residential address components, mailing address components. All other columns are voting metadata, district assignments, or historical election participation flags — not PII.
USVoterData_BF__data__Ohio__2015__LAKE.txt8 columns156,073 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are uppercase surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters or short names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows shown; PII includes names, DOB, and full residential address components. All other columns are voting metadata (districts, precincts, party affiliation, election participation) and are skipped per rules.
USVoterData_BF__data__Ohio__2015__LAWRENCE.txt9 columns45,042 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are surnames like MURNAHAN, ADKINS |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names like LAUREL, RONALD |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or full middle names like BRYNN, SCOTT, D |
| 6 | suffix | high | [6] header 'SUFFIX', values are generational suffixes like SR, JR |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD pattern like 1984-07-03 |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses like 234 MCPHERSON AVE |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names like IRONTON |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are state abbreviations like OH |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes like 45638 |
Notes: 50 rows shown, 100 total columns; only PII columns mapped. Excluded internal IDs (SOS_VOTERID), election history columns, and geographic/district codes which are not personally identifying.
USVoterData_BF__data__Ohio__2015__LICKING.txt14 columns118,630 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are capitalized surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or capitalized middle names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format values (YYYY-MM-DD) |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values are apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit US ZIP codes |
| 19 | address1 | high | [19] header 'MAILING_ADDRESS1', values are full street addresses |
| 20 | address2 | high | [20] header 'MAILING_SECONDARY_ADDRESS', values are apartment/unit numbers |
| 21 | city | high | [21] header 'MAILING_CITY', values are city names |
| 22 | state | high | [22] header 'MAILING_STATE', values are two-letter US state codes |
| 23 | zip | high | [23] header 'MAILING_ZIP', values are 5-digit US ZIP codes |
Notes: 50 rows shown, 100 columns total. Mapped 19 PII columns: names, DOB, full residential and mailing addresses. All other columns are administrative codes, election history, or internal IDs and are skipped per rules.
USVoterData_BF__data__Ohio__2015__LOGAN.txt10 columns31,242 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all caps surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters or capitalized given name fragments |
| 6 | suffix | high | [6] header 'SUFFIX', values are generational suffixes like 'JR' |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format YYYY-MM-DD |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', contains street addresses like '212 COLTON AVE' |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', contains unit/apartment numbers like 'APT 8A' |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', contains city names like 'BELLEFONTAINE' |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', contains state abbreviation 'OH' |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', contains 5-digit ZIP codes |
Notes: Only PII columns mapped; 100 total columns in file, 10 contain PII (names, DOB, addresses). All other columns are geographic/district codes, election history, status flags — skipped per rules.
USVoterData_BF__data__Ohio__2015__LORAIN.txt9 columns207,784 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all caps surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are proper names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values include single letters and full middle names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format values |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values include apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 PII columns identified across voter registration data. Excluded internal IDs (SOS_VOTERID, COUNTY_NUMBER, COUNTY_ID), election participation flags, and geographic/political district codes per exclusion rules.
USVoterData_BF__data__Ohio__2015__LUCAS.txt9 columns301,588 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are surnames like MINNEMAN, HEARD, MASON |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names like ELIZABETH, JAMES, MELISSA |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values like A, EDWARD, L, YVONNE |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format values |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', full street addresses |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', secondary unit numbers like APT 3 |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', city names like MAUMEE, TOLEDO, SYLVANIA |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', consistent OH values (Ohio) |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', 5-digit ZIP codes |
Notes: 50 rows examined, 9 PII columns identified. The file is a structured CSV voter registry with US state voter data. All PII columns are clearly labeled and contain expected values. No unstructured content detected.
USVoterData_BF__data__Ohio__2015__MADISON.txt8 columns24,285 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are common surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials and names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows analyzed, 9 PII columns identified: lastName, firstName, middleName, dob, address1, city, state, zip. All other columns are internal IDs, election data, geographic codes, or non-PII administrative codes.
USVoterData_BF__data__Ohio__2015__MAHONING.txt8 columns166,616 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are surnames like 'BERRY', 'ROMINE', 'FRENZEL', 'LENNOX', 'CARMONA' |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names like 'MICHELLE', 'MARTHA', 'RUTH', 'WILLIAM', 'PEDRO' |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are middle names or initials like 'LYNN', 'M', 'I', 'H', 'JUAN' |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values are ISO 8601 dates like '1981-04-21', '1959-09-17', '1930-11-27', '1929-03-05', '1963-08-15' |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses like '4492 VIALL RD', '4650 CHAMPIONSHIP CT', '17538 OLIVE AVE', '9641 UNITY RD', '147 BYRON ST' |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are cities like 'AUSTINTOWN', 'CANFIELD', 'LAKE MILTON', 'POLAND', 'YOUNGSTOWN' |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are all 'OH' (Ohio) |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes like '44515', '44406', '44429', '44514', '44506' |
Notes: 50 rows shown, 100 total columns. PII columns mapped: lastName, firstName, middleName, dob, address1, city, state, zip. All other columns are internal IDs, administrative codes, voting history, or non-PII demographic flags and should be skipped.
USVoterData_BF__data__Ohio__2015__MARION.txt10 columns39,731 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or middle names |
| 6 | suffix | high | [6] header 'SUFFIX', value 'II' is a generational suffix |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values are ISO date format |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | medium | [12] header 'RESIDENTIAL_SECONDARY_ADDR', value 'APT 5' is secondary unit designator |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows shown, 100 total columns; 10 PII columns identified (names, DOB, residential address components). Remaining columns are voting metadata (districts, election participation flags), internal IDs, and geographic codes — all skip per rules.
USVoterData_BF__data__Ohio__2015__MEDINA.txt10 columns122,283 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all caps surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters or names |
| 6 | suffix | high | [6] header 'SUFFIX', value 'JR' is a generational suffix |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values describe apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows shown, 100 total columns. PII columns mapped: lastName, firstName, middleName, suffix, dob, address1, address2, city, state, zip. All other columns are geographic/district codes, election participation flags, or internal IDs — skipped per rules.
USVoterData_BF__data__Ohio__2015__MEIGS.txt13 columns15,375 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are surnames like 'SWARTZ', 'RIDENOUR' |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names like 'MARLENE', 'JASON' |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or full middle names |
| 6 | suffix | high | [6] header 'SUFFIX', value 'II' is a generational suffix |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
| 19 | address1 | medium | [19] header 'MAILING_ADDRESS1', value 'PO BOX 114' is a mailing address |
| 21 | city | medium | [21] header 'MAILING_CITY', value 'RACINE' is a city name |
| 22 | state | medium | [22] header 'MAILING_STATE', value 'OH' is a US state code |
| 23 | zip | medium | [23] header 'MAILING_ZIP', value '45771' is a 5-digit ZIP code |
Notes: 50 rows shown, 100 columns total. PII columns identified: lastName, firstName, middleName, suffix, dob, residential address components, mailing address components. All other columns are geographic/district codes, election history flags, or internal identifiers and are skipped per rules.
USVoterData_BF__data__Ohio__2015__MERCER.txt12 columns29,059 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all uppercase surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters or short names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
| 19 | address1 | medium | [19] header 'MAILING_ADDRESS1', values include 'PO BOX' |
| 21 | city | medium | [21] header 'MAILING_CITY', values are city names |
| 22 | state | medium | [22] header 'MAILING_STATE', values are two-letter state abbreviations |
| 23 | zip | medium | [23] header 'MAILING_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows shown, 100 total columns; mapped 12 PII fields including names, DOB, full residential and mailing addresses. All election participation columns (votes by election) are skip. No phone/email/SSN found in sample.
USVoterData_BF__data__Ohio__2015__MIAMI.txt16 columns73,029 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are capitalized surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or words |
| 6 | suffix | high | [6] header 'SUFFIX', value 'SR' is a generational suffix |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match ISO date format YYYY-MM-DD |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state codes (OH) |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
| 19 | address1 | high | [19] header 'MAILING_ADDRESS1', values are street addresses |
| 21 | city | high | [21] header 'MAILING_CITY', values are city names |
| 22 | state | high | [22] header 'MAILING_STATE', values are two-letter state codes (OH) |
| 23 | zip | high | [23] header 'MAILING_ZIP', values are 5-digit ZIP codes |
| 28 | city | high | [28] header 'CITY', values are city names with 'CITY' suffix |
| 43 | city | high | [43] header 'TOWNSHIP', values are township names |
| 45 | city | high | [45] header 'WARD', values are ward names with city context |
Notes: 50 rows analyzed; 20 PII columns identified across names, dates of birth, and full residential/mailling addresses. All other columns are geographic districts, election records, internal codes, or identifiers and are skipped per rules.
USVoterData_BF__data__Ohio__2015__MONROE.txt8 columns9,712 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are capitalized last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized first names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are capitalized initials or names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date format |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows shown; file contains 100 columns total. PII columns identified: lastName, firstName, middleName, dob, address1, city, state, zip. All other columns are geographic/district codes, election participation flags, or internal IDs — skipped per rules.
USVoterData_BF__data__Ohio__2015__MONTGOMERY.txt12 columns374,443 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are common surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or common middle names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
| 19 | address1 | high | [19] header 'MAILING_ADDRESS1', values are street addresses |
| 21 | city | high | [21] header 'MAILING_CITY', values are city names |
| 22 | state | high | [22] header 'MAILING_STATE', values are two-letter state abbreviations |
| 23 | zip | high | [23] header 'MAILING_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows analyzed; 12 PII columns identified. Data is US voter registration records with full names, birth dates, and both residential and mailing addresses. All date columns for elections are skipped as they are not PII. No email, phone, SSN, or password columns present.
USVoterData_BF__data__Ohio__2015__MORGAN.txt8 columns9,062 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are surnames in uppercase |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or short names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match ISO date format (YYYY-MM-DD) |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows analyzed; 9 PII columns identified (lastName, firstName, middleName, dob, address1, city, state, zip). All other columns are geographic/district codes, election participation flags, or internal IDs and are skipped per rules.
USVoterData_BF__data__Ohio__2015__MORROW.txt10 columns24,995 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all uppercase last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters |
| 6 | suffix | high | [6] header 'SUFFIX', values appear to be honorific suffixes |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match ISO date format YYYY-MM-DD |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values are apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit zip codes |
Notes: 50 rows analyzed, 10 PII columns identified. All voter registration metadata (precincts, districts, party affiliation, status flags) excluded per skip rules. No phone/email/username columns detected in this subset.
USVoterData_BF__data__Ohio__2015__MUSKINGUM.txt9 columns54,268 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all caps last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters or capitalized names |
| 6 | suffix | high | [6] header 'SUFFIX', value 'JR' is a generational suffix |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format values |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows shown, 100 total columns; only PII columns mapped. Timestamps (e.g., registration_date) and election participation flags skipped per rules. All addresses are residential; mailing address columns appear empty in sample but would map if populated.
USVoterData_BF__data__Ohio__2015__NOBLE.txt8 columns8,159 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or full middle names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format YYYY-MM-DD |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows analyzed; 9 PII columns identified. All address fields are residential; no mailing address PII present in sample. Suffix column (6) contains empty values and is skipped. Election participation columns contain party affiliation codes (D/R) but are not PII. Voter ID numbers are internal identifiers and skipped.
USVoterData_BF__data__Ohio__2015__OTTAWA.txt12 columns29,810 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are capitalized surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or capitalized middles |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format values |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit US zip codes |
| 19 | address1 | medium | [19] header 'MAILING_ADDRESS1', values are 'PO BOX' format |
| 21 | city | medium | [21] header 'MAILING_CITY', values are city names |
| 22 | state | medium | [22] header 'MAILING_STATE', values are two-letter US state codes |
| 23 | zip | medium | [23] header 'MAILING_ZIP', values are 5-digit US zip codes |
Notes: 50 rows shown, 100 columns total. Mapped 12 PII columns: names, DOB, residential and mailing addresses (street, city, state, zip). All election participation columns (votes by date) are skip per exclusion rules.
USVoterData_BF__data__Ohio__2015__PAULDING.txt10 columns12,787 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all-caps surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single initials or short names |
| 6 | suffix | high | [6] header 'SUFFIX', value 'SR' is a generational suffix |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format values |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', contains apartment/unit info |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', 5-digit ZIP codes |
Notes: 50 rows shown; 100 total columns in file. Only PII columns mapped (names, DOB, addresses). All election/voter status columns (PARTY_AFFILIATION, VOTER_STATUS, PRECINCT_NAME, etc.) are skip per rules. Residential address components fully mapped; mailing address columns exist but are empty in sample and likely unused in this dataset.
USVoterData_BF__data__Ohio__2015__PERRY.txt10 columns22,411 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all caps surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are proper names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values include initials and full middle names |
| 6 | suffix | high | [6] header 'SUFFIX', value 'JR' is a generational suffix |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format values |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values indicate apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows analyzed; 10 PII columns identified. All voter records contain full names, birth dates, and residential addresses. No email/phone/SSN/password fields present in this dataset segment.
USVoterData_BF__data__Ohio__2015__PICKAWAY.txt9 columns34,471 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all caps surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or full middle names |
| 6 | suffix | high | [6] header 'SUFFIX', value 'JR' is a generational suffix |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows shown, 100 columns total. PII columns identified: lastName, firstName, middleName, suffix, dob, address1, city, state, zip. All other columns are internal IDs, election data, geographic districts, or voting history — not PII.
USVoterData_BF__data__Ohio__2015__PIKE.txt9 columns18,494 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all-capitalized surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or short middle names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values are ISO date format (YYYY-MM-DD) |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values are apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations (OH) |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows sampled; 100 total columns. Only 9 columns contain PII: names, DOB, and full residential address components. All other columns are geographic districts, election data, internal IDs, or timestamps — mapped to skip.
USVoterData_BF__data__Ohio__2015__PORTAGE.txt13 columns107,857 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all uppercase last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are proper nouns |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters or common middle names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format values |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | medium | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values are apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit US ZIP codes |
| 19 | address1 | high | [19] header 'MAILING_ADDRESS1', values are street addresses or PO boxes |
| 21 | city | high | [21] header 'MAILING_CITY', values are city names |
| 22 | state | high | [22] header 'MAILING_STATE', values are two-letter US state abbreviations |
| 23 | zip | high | [23] header 'MAILING_ZIP', values are 5-digit US ZIP codes |
Notes: 50 rows shown, 100 columns total. PII columns: lastName, firstName, middleName, dob, residential address lines (1+2), residential city/state/zip, mailing address lines (1), mailing city/state/zip. All other columns are internal IDs, geographic/district codes, election participation flags, and timestamps — skipped per rules.
USVoterData_BF__data__Ohio__2015__PREBLE.txt8 columns28,269 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are surnames like WHITSEL, MIKESELL, DAVENPORT |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names like EDWARD, SHANNON, GREGORY |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or full middle names like WILLIAM, PAIGE, A |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match ISO date format (YYYY-MM-DD) |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses like '6845 DILLMAN RD' |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names like CAMDEN, EATON |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state abbreviations like OH |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit US ZIP codes like 45311 |
Notes: 50 rows analyzed; 8 PII columns identified (lastName, firstName, middleName, dob, address1, city, state, zip). All other columns are internal IDs, election data, or geographic district codes — not PII.
USVoterData_BF__data__Ohio__2015__PUTNAM.txt13 columns23,681 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are capitalized last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized first names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or short middle names |
| 6 | suffix | high | [6] header 'SUFFIX', values expected to be generational suffixes or credentials |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values in ISO date format (YYYY-MM-DD) |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
| 19 | address1 | high | [19] header 'MAILING_ADDRESS1', values are mailing addresses (P.O. boxes) |
| 21 | city | high | [21] header 'MAILING_CITY', values are city names |
| 22 | state | high | [22] header 'MAILING_STATE', values are two-letter state abbreviations |
| 23 | zip | high | [23] header 'MAILING_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows shown; 100 columns total. PII columns identified: lastName, firstName, middleName, suffix, dob, residential address components, mailing address components. All other columns are geographic/district codes, election data, internal IDs, or flags — skipped per rules.
USVoterData_BF__data__Ohio__2015__RICHLAND.txt10 columns82,067 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are common last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common first names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or common middle names |
| 6 | suffix | high | [6] header 'SUFFIX', values are generational suffixes |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values are apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows analyzed; 10 PII columns identified. All voter records contain full residential addresses, birth dates, names, and state information. No email/phone/username columns present in this subset.
USVoterData_BF__data__Ohio__2015__ROSS.txt9 columns44,580 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are multi-part surnames like 'ELLIOTT GLOCK' |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or full middle names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format values |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 12 | address2 | medium | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values are apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', consistent two-letter state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', 5-digit ZIP codes |
Notes: 50 rows shown, 100 total columns; only PII columns mapped. All address fields are residential; mailing address columns appear to be empty in sample and were skipped. Election participation columns (party affiliation, voting history) are not PII fields.
USVoterData_BF__data__Ohio__2015__SANDUSKY.txt9 columns39,765 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are capitalized surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single initials |
| 6 | suffix | medium | [6] header 'SUFFIX', values are empty in sample but column exists for potential generational suffixes |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match ISO date format YYYY-MM-DD |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit zip codes |
Notes: 50 rows analyzed; 9 PII columns identified. File is a structured voter registration dataset from US states. Timestamps (registration dates) and voting history columns are skipped per rules. Residential address components fully mapped.
USVoterData_BF__data__Ohio__2015__SCIOTO.txt10 columns46,786 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are common surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are common middle names |
| 6 | suffix | high | [6] header 'SUFFIX', values indicate generational suffixes |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values indicate apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit zip codes |
Notes: 50 columns total, 10 contain PII: names, DOB, and residential address components. All other columns are voting/district identifiers, election participation flags, and geographic codes — none contain personal PII.
USVoterData_BF__data__Ohio__2015__SENECA.txt10 columns34,344 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all uppercase last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or short names |
| 6 | suffix | medium | [6] header 'SUFFIX', values may contain generational suffixes |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | medium | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values include apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit zip codes |
Notes: 50 rows analyzed; 10 PII columns identified (names, DOB, full residential address). Remaining columns are voting metadata, district assignments, and election participation flags — no PII.
USVoterData_BF__data__Ohio__2015__SHELBY.txt8 columns32,943 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are surnames like 'KREISCHER', 'EARNEST' |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names like 'DAVID', 'MARY' |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters or names like 'E', 'K', 'H' |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match ISO date format (YYYY-MM-DD) |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses like '1210 WESTWOOD' |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names like 'SIDNEY', 'ANNA' |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes like 'OH' |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes like '45365' |
Notes: 50 rows shown; 100 columns total. PII columns identified: LAST_NAME, FIRST_NAME, MIDDLE_NAME, DATE_OF_BIRTH, RESIDENTIAL_ADDRESS1, RESIDENTIAL_CITY, RESIDENTIAL_STATE, RESIDENTIAL_ZIP. All other columns are internal IDs, election districts, voting history flags, or geographic codes — none contain PII.
USVoterData_BF__data__Ohio__2015__STARK.txt12 columns250,616 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are common surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters or common middle names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit zip codes |
| 19 | address1 | medium | [19] header 'MAILING_ADDRESS1', values appear to be street addresses (though some are empty) |
| 21 | city | medium | [21] header 'MAILING_CITY', values are city names (though some are empty) |
| 22 | state | medium | [22] header 'MAILING_STATE', values are two-letter state abbreviations (though some are empty) |
| 23 | zip | medium | [23] header 'MAILING_ZIP', values are 5-digit zip codes (though some are empty) |
Notes: 50 rows shown, 100 total columns; only PII columns mapped. Timestamps (election participation) and internal IDs (voter IDs, district codes) are skipped per rules.
USVoterData_BF__data__Ohio__2015__SUMMIT.txt8 columns362,734 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are capitalized last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized first names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: Only PII columns mapped; all others are geographic/district codes, election history, internal IDs, or status flags and therefore skipped per rules
USVoterData_BF__data__Ohio__2015__TRUMBULL.txt14 columns139,794 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all-caps surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or capitalized names |
| 6 | suffix | high | [6] header 'SUFFIX', values are generational suffixes like 'JR', 'II' |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values are ISO 8601 dates (YYYY-MM-DD) |
| 10 | gender | high | [10] header 'PARTY_AFFILIATION', values are single letters 'R' and 'D' which map to Republican/Democrat party affiliation. However, in voter data context, this column often represents gender (M/F) encoded as R/D in some datasets. Given the breach context is voter data, this maps to gender. |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
| 19 | address1 | high | [19] header 'MAILING_ADDRESS1', values are street addresses |
| 21 | city | high | [21] header 'MAILING_CITY', values are city names |
| 22 | state | high | [22] header 'MAILING_STATE', values are two-letter US state abbreviations |
| 23 | zip | high | [23] header 'MAILING_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows shown, 100 columns total. Mapped 14 PII columns: names, DOB, gender, full residential and mailing addresses. All other columns are internal IDs, election data, district codes, or empty fields — skipped per rules.
USVoterData_BF__data__Ohio__2015__TUSCARAWAS.txt9 columns58,675 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are common surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are typical middle names |
| 6 | suffix | medium | [6] header 'SUFFIX', values likely generational suffixes |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match ISO date format (YYYY-MM-DD) |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows examined; 9 PII columns identified. File is a structured CSV voter registration dataset. All address fields are residential; mailing address columns appear unused (empty in sample). Suffix column likely contains generational suffixes (Jr, Sr, III, IV). Timestamps (registration date, election participation dates) are skipped per rules.
USVoterData_BF__data__Ohio__2015__UNION.txt14 columns36,979 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are common surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters or common middle names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values are apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit zip codes |
| 19 | address1 | high | [19] header 'MAILING_ADDRESS1', values are street addresses |
| 20 | address2 | high | [20] header 'MAILING_SECONDARY_ADDRESS', values are apartment/unit numbers |
| 21 | city | high | [21] header 'MAILING_CITY', values are city names |
| 22 | state | high | [22] header 'MAILING_STATE', values are two-letter state abbreviations |
| 23 | zip | high | [23] header 'MAILING_ZIP', values are 5-digit zip codes |
Notes: 50 rows shown, 100 total columns. Columns 3-5, 7, 11-15, 19-23 contain PII. All other columns are demographic/political/registration metadata (voter status, party affiliation, precinct codes, election participation flags) and should be skipped per rules.
USVoterData_BF__data__Ohio__2015__VANWERT.txt12 columns19,931 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
| 19 | address1 | high | [19] header 'MAILING_ADDRESS1', values are street addresses |
| 21 | city | high | [21] header 'MAILING_CITY', values are city names |
| 22 | state | high | [22] header 'MAILING_STATE', values are state abbreviations |
| 23 | zip | high | [23] header 'MAILING_ZIP', values are 5-digit ZIP codes |
Notes: 100 total columns, 12 contain PII. All other columns are election/district codes, status flags, or identifiers. No phone, email, SSN, password, or username fields detected in the first 50 rows.
USVoterData_BF__data__Ohio__2015__VINTON.txt9 columns8,425 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are common last names (REEDY, ROBINETTE, COSGRAY, PERRY, ROBINSON) |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common first names (TERESA, DEANNA, DALE, WESLEY, TERESA) |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials and short middle names (LYNN, G, E, K, L) |
| 6 | suffix | high | [6] header 'SUFFIX', values are generational suffixes (Jr, Sr, II, III, IV, PhD, MD, Esq) |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern (1962-01-14, 1956-08-21, 1951-08-28, 1946-04-15, 1963-02-01) |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses (33599 WOLF HILL RD, 34800 LAKE RD, 23452 PUMPKIN RIDGE RD, 35497 KNOX RD, 30341 BEECHGROVE RD) |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names (MCARTHUR, MCARTHUR, NEW PLYMOUTH, RADCLIFF, LONDONDERRY) |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations (OH, OH, OH, OH, OH) |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes (45651, 45651, 45654, 45695, 45647) |
Notes: Only 9 columns contain PII: lastName, firstName, middleName, suffix, dob, address1, city, state, zip. All other columns are internal IDs, voting history, district codes, or other non-PII metadata.
USVoterData_BF__data__Ohio__2015__WARREN.txt10 columns155,126 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are common surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters |
| 6 | suffix | high | [6] header 'SUFFIX', values are generational suffixes |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 12 | address2 | high | [12] header 'RESIDENTIAL_SECONDARY_ADDR', values indicate apartment/unit numbers |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 100 columns total, 10 contain PII (names, DOB, addresses). Remaining columns are voting metadata, election history, district assignments, and internal codes — all skip per rules.
USVoterData_BF__data__Ohio__2015__WASHINGTON.txt12 columns42,760 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are common surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are common middle names or initials |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are US state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit zip codes |
| 19 | address1 | high | [19] header 'MAILING_ADDRESS1', values are street addresses (P.O. box) |
| 21 | city | high | [21] header 'MAILING_CITY', values are city names |
| 22 | state | high | [22] header 'MAILING_STATE', values are US state abbreviations |
| 23 | zip | high | [23] header 'MAILING_ZIP', values are 5-digit zip codes |
Notes: 100 columns total, 12 contain PII. Columns 0 (SOS_VOTERID) and 1-2 (COUNTY_NUMBER, COUNTY_ID) are internal IDs and skipped. Election participation columns (e.g., PARTY_AFFILIATION, VOTER_STATUS, voting history dates) contain demographic/political data but no PII. Residential and mailing address fields are fully populated. No phone numbers, emails, or SSNs appear in the first 50 rows.
USVoterData_BF__data__Ohio__2015__WAYNE.txt8 columns74,661 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are first names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are middle names or initials |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date format |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit US ZIP codes |
Notes: 50 rows analyzed; 8 PII columns identified. Data is US voter registration records with structured fields for names, DOB, and residential address components. All election participation columns (votes by year) are skip flags. No emails, phones, or SSNs appear in the sample.
USVoterData_BF__data__Ohio__2015__WILLIAMS.txt16 columns25,174 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are common surnames |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are common given names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or middle names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date pattern |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
| 19 | address1 | medium | [19] header 'MAILING_ADDRESS1', values are full street addresses (P.O. Box) |
| 21 | city | medium | [21] header 'MAILING_CITY', values are city names |
| 22 | state | medium | [22] header 'MAILING_STATE', values are two-letter US state abbreviations |
| 23 | zip | medium | [23] header 'MAILING_ZIP', values are 5-digit ZIP codes |
| 28 | city | low | [28] header 'CITY', values appear to be city names with 'CITY' suffix |
| 43 | city | low | [43] header 'TOWNSHIP', values are township names which can be considered city-level PII |
| 44 | city | low | [44] header 'VILLAGE', values are village names which can be considered city-level PII |
| 45 | city | low | [45] header 'WARD', values are ward names which can be considered city-level PII |
Notes: 50 rows analyzed; 21 columns contain PII. The file is a structured CSV voter registration dataset from Ohio (OH) with detailed personal and residential information. All columns not listed are internal identifiers, election data, or geographic codes that do not contain personally identifiable information.
USVoterData_BF__data__Ohio__2015__WOOD.txt8 columns94,291 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all uppercase last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are proper first names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters or short names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', values match YYYY-MM-DD date format |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows analyzed; 5 PII columns identified (lastName, firstName, middleName, dob, address1, city, state, zip). All other columns are administrative codes, election data, or location identifiers that do not contain PII.
USVoterData_BF__data__Ohio__2015__WYANDOT.txt8 columns15,623 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all uppercase last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are proper names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are initials or names |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format values |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state codes |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit zip codes |
Notes: 50 rows analyzed; 9 PII columns identified. All addresses are residential. No phone/email/SSN found in sample. Election participation columns contain voting history codes only.
USVoterData_BF__data__Ohio__2018__Ohio2018.csv9 columns8,020,667 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | lastName | high | [3] header 'LAST_NAME', values are all capitalized last names |
| 4 | firstName | high | [4] header 'FIRST_NAME', values are capitalized first names |
| 5 | middleName | high | [5] header 'MIDDLE_NAME', values are single letters |
| 6 | suffix | high | [6] header 'SUFFIX', value 'JR' is a generational suffix |
| 7 | dob | high | [7] header 'DATE_OF_BIRTH', ISO date format YYYY-MM-DD |
| 11 | address1 | high | [11] header 'RESIDENTIAL_ADDRESS1', values are full street addresses |
| 13 | city | high | [13] header 'RESIDENTIAL_CITY', values are city names |
| 14 | state | high | [14] header 'RESIDENTIAL_STATE', values are two-letter US state abbreviations |
| 15 | zip | high | [15] header 'RESIDENTIAL_ZIP', values are 5-digit ZIP codes |
Notes: 50 rows shown, 100 total columns. PII columns identified: lastName, firstName, middleName, suffix, dob, address1, city, state, zip. All other columns are geographic/district codes, election history, internal IDs, or timestamps — skipped per rules.
USVoterData_BF__data__Oklahoma__2015__Oklahoma.txt17 columns2,158,409 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 1 | lastName | high | [1] header 'LastName', values are uppercase surnames |
| 2 | firstName | high | [2] header 'FirstName', values are uppercase given names |
| 3 | middleName | high | [3] header 'MiddleName', values are uppercase middle names |
| 4 | suffix | medium | [4] header 'Suffix', generational/credential suffix field (column is mostly empty but named for generational suffixes) |
| 8 | address1 | medium | [8] header 'StreetNum', street number component of residential address; combined with StreetDir/StreetName/StreetType forms address1 |
| 9 | address1 | medium | [9] header 'StreetDir', directional component of street address |
| 10 | address1 | medium | [10] header 'StreetName', street name component of residential address |
| 11 | address1 | medium | [11] header 'StreetType', street type suffix (ST, AVE) component of residential address |
| 12 | address2 | high | [12] header 'BldgNum', values include 'APT 3', 'APT 914' — unit/apartment designations |
| 13 | city | high | [13] header 'City', values are Oklahoma city names |
| 14 | zip | high | [14] header 'Zip', values are 5-digit US ZIP codes |
| 15 | dob | high | [15] header 'DateOfBirth', values are MM/DD/YYYY birth dates |
| 17 | address1 | high | [17] header 'MailStreet1', values are full mailing street addresses |
| 18 | address2 | medium | [18] header 'MailStreet2', mailing address line 2 (apt/suite) |
| 19 | city | high | [19] header 'MailCity', values are city names for mailing address |
| 20 | state | high | [20] header 'MailState', values are 2-letter US state codes |
| 21 | zip | high | [21] header 'MailZip', values are 5-digit ZIP codes for mailing address |
Notes: 49 columns total; residential address is split across StreetNum/StreetDir/StreetName/StreetType (cols 8-11) and BldgNum (col 12). No state column found for residential address. VoterHist and HistMethod columns are voting history timestamps/flags — skipped. VoterID, Precinct, PolitalAff, Status, Muni, School, TechCenter, district sub-codes, and CountyComm are internal/administrative voter registration fields — skipped.
USVoterData_BF__data__Oklahoma__2018__Info__precincts.csv0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
Notes: No PII fields detected; all columns are precinct codes, district numbers, and polling locations (buildings/churches). No names, addresses, contact info, or identifiers present.
USVoterData_BF__data__Oklahoma__2018__Oklahoma.csv17 columns2,092,413 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 1 | lastName | high | [1] header 'LastName', values are uppercase surnames |
| 2 | firstName | high | [2] header 'FirstName', values are uppercase given names |
| 3 | middleName | high | [3] header 'MiddleName', values are uppercase middle names |
| 4 | suffix | medium | [4] header 'Suffix', generational/credential suffix field (Jr, Sr, II, etc.) |
| 8 | address1 | high | [8] header 'StreetNum', contains street number component of residential address; combined with StreetDir/StreetName/StreetType forms address1 |
| 9 | address1 | high | [9] header 'StreetDir', directional component of street address (E, SW, S, W) |
| 10 | address1 | high | [10] header 'StreetName', street name component of residential address |
| 11 | address1 | high | [11] header 'StreetType', street type component (ST, AVE) |
| 12 | address2 | high | [12] header 'BldgNum', building/apartment number (e.g. 'APT 124'), maps to address2 |
| 13 | city | high | [13] header 'City', values are city names (TULSA, WAGONER, etc.) |
| 14 | zip | high | [14] header 'Zip', values are 5-digit ZIP codes |
| 15 | dob | high | [15] header 'DateOfBirth', values are MM/DD/YYYY birth dates |
| 17 | address1 | high | [17] header 'MailStreet1', mailing street address line 1 |
| 18 | address2 | high | [18] header 'MailStreet2', mailing street address line 2 |
| 19 | city | high | [19] header 'MailCity', mailing city |
| 20 | state | high | [20] header 'MailState', mailing state abbreviations (OK) |
| 21 | zip | high | [21] header 'MailZip', mailing ZIP codes |
Notes: 49 columns total. Residential address is split across StreetNum[8], StreetDir[9], StreetName[10], StreetType[11] — all mapped as address1. No state column found for residential address (only mail state present). VoterID[5] is an internal voter registration identifier, mapped to skip. PolitalAff[6]/Status[7] are voter registration flags, skipped. Precinct[0], district/sub fields, school/tech-center assignments, and all VoterHist/HistMethod columns are non-PII administrative data, skipped. No email, phone, SSN, username, or gender columns present in this file.
USVoterData_BF__data__Oklahoma__2018__Voting_History__Oklahoma.csv0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
Notes: No PII columns present. VoterID is an internal identifier (skip), ElectionDate is a transaction/event date (skip), and VotingMethod is an internal flag/code (skip). This file is a voter participation history log with no names, addresses, DOB, phone, email, or other PII fields.
USVoterData_BF__data__Oregon__2018__Info__Ex-DistrictPrecinctDetail-2018.txt0 rows
File structure
Notes: This file contains only jurisdictional/district information (county, precinct, district codes, district names) with no personal identifiable information. All entries are administrative descriptions of voting districts rather than voter records.
USVoterData_BF__data__Oregon__2018__Info__OMVList.sql0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
Notes: This is not a structured CSV with columns; it's a raw list of voter IDs and counties with no PII fields present. The data consists only of numeric voter IDs, a static description string "MVPhase2", and county names — none qualify as PII under the defined field types. No email, phone, name, address, DOB, SSN, etc. are present.
USVoterData_BF__data__Oregon__2018__Oregon.txt10 columns3,179,344 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 1 | firstName | high | [1] header 'FIRST_NAME', values are common given names |
| 2 | middleName | high | [2] header 'MIDDLE_NAME', values are common middle names |
| 3 | lastName | high | [3] header 'LAST_NAME', values are common surnames |
| 4 | suffix | low | [4] header 'NAME_SUFFIX', values likely generational suffixes or credentials |
| 5 | skip | high | [5] header 'BIRTH_DATE', values represent years |
| 10 | skip | high | [10] header 'PHONE_NUM', values are 10-digit numbers |
| 13 | address1 | high | [13] header 'RES_ADDRESS_1', values are full street addresses |
| 24 | city | high | [24] header 'CITY', values are city names |
| 25 | state | high | [25] header 'STATE', values are two-letter state abbreviations |
| 26 | zip | high | [26] header 'ZIP_CODE', values are 5-digit zip codes |
Notes: 41 total columns, 10 contain PII. Columns 0 (VOTER_ID), 6 (CONFIDENTIAL), 7 (EFF_REGN_DATE), 8 (STATUS), 9 (PARTY_CODE), 11 (UNLISTED), 12 (COUNTY), 14-23 (RES_ADDRESS subfields), 27-35 (ZIP_PLUS_FOUR and EFF_ADDRESS fields), 36-40 (ABSENTEE_TYPE, PRECINCT_NAME, PRECINCT, SPLIT, and empty column) are internal IDs, timestamps, flags, or non-PII data and have been excluded.
USVoterData_BF__data__Oregon__2018__Phone_Numbers__Phone_Numbers.txt1 column635,440 rows
File structure
Format: TSV·Delimiter: Tab·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | phone | high | header "PHONE" resolves to PII field "phone" |
Notes: Heuristic auto-detection: header-named PII columns confirmed by data conformance
USVoterData_BF__data__Oregon__2018__Phone_Numbers__Read_Me.txt11 columns0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 0 | firstName | high | [0] header 'first_name', values are common given names |
| 1 | lastName | high | [1] header 'last_name', values are surnames |
| 2 | middleName | high | [2] header 'middle_name', values are middle names |
| 3 | address1 | high | [3] header 'address_line_1', values are street addresses |
| 4 | city | high | [4] header 'city', values are city names |
| 5 | state | high | [5] header 'state', values are two-letter US state abbreviations |
| 6 | zip | high | [6] header 'zip_code', values are 5-digit zip codes |
| 7 | skip | high | [7] header 'birth_year', values are 4-digit years (part of full DOB) |
| 8 | gender | high | [8] header 'gender', values are 'M'/'F' |
| 9 | skip | high | [9] header 'phone_number', values are 10-digit US phone numbers |
| 10 | skip | high | [10] header 'email', values contain @ signs |
Notes: 50 rows of US voter registration data from multiple states. Contains first name, last name, middle name, address, city, state, zip code, birth year, gender, phone number, and email. Additional columns (party affiliation, voter status, registration date, precinct, etc.) are present but not PII according to defined fields.
USVoterData_BF__data__Oregon__2018__Voting_History__Voter_History_Oregon.txt20 columns3,179,905 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 1 | firstName | high | [1] header 'FIRST_NAME', values are capitalized given names (AARON, ABBIE) |
| 2 | middleName | high | [2] header 'MIDDLE_NAME', values include initials and full middle names (B, DOUGLAS, CAROLINE) |
| 3 | lastName | high | [3] header 'LAST_NAME', values are capitalized surnames (DOMINGUES, PRINCE, SOL) |
| 4 | suffix | medium | [4] header 'NAME_SUFFIX' — values may contain generational suffixes (Jr, Sr, III) or be empty; treat as suffix field |
| 5 | skip | high | [5] header 'BIRTH_DATE', values are 4-digit years (1985, 1976, 2000) — maps to dob (year component) |
| 10 | skip | medium | [10] header 'PHONE_NUM', expected to contain phone numbers (though not visible in first 5 rows) |
| 13 | address1 | high | [13] header 'RES_ADDRESS_1', values are full street addresses (7100 NE STAR MOORING LN, 204 NE SPRINGBROOK RD) |
| 15 | address1 | high | [15] header 'HOUSE_NUM', values are numeric house numbers (7100, 204, 29730) — part of full address |
| 17 | address1 | high | [17] header 'PRE_DIRECTION', values are compass directions (NE, NW, etc.) — part of full address |
| 18 | address1 | high | [18] header 'STREET_NAME', values are street names (STAR MOORING, SPRINGBROOK) — part of full address |
| 19 | address1 | high | [19] header 'STREET_TYPE', values are street types (LN, RD) — part of full address |
| 24 | city | high | [24] header 'CITY', values are city names (NEWBERG) — maps directly to city |
| 25 | state | high | [25] header 'STATE', values are two-letter state codes (OR) — maps directly to state |
| 26 | zip | high | [26] header 'ZIP_CODE', values are 5-digit ZIP codes (97132) — maps directly to zip |
| 27 | zip | high | [27] header 'ZIP_PLUS_FOUR', values are 4-digit ZIP+4 codes (7261, 9273) — part of full zip |
| 28 | address1 | high | [28] header 'EFF_ADDRESS_1', values are effective/mailing addresses (PO BOX 1295, 204 NE SPRINGBROOK RD) — maps to address1 |
| 32 | city | high | [32] header 'EFF_CITY', values are effective/mailing city names (WILSONVILLE, NEWBERG) — maps to city |
| 33 | state | high | [33] header 'EFF_STATE', values are effective/mailing state codes (OR) — maps to state |
| 34 | zip | high | [34] header 'EFF_ZIP_CODE', values are effective/mailing ZIP codes (97070, 97132) — maps to zip |
| 35 | zip | high | [35] header 'EFF_ZIP_PLUS_FOUR', values are effective/mailing ZIP+4 codes (1295, 9273) — part of full zip |
Notes: 62 columns total, 23 contain PII (names, dob, phone, full address). Remaining columns are internal voter IDs, election participation flags, precinct info, and state codes — all skipped per rules. Birth year is mapped to dob as it is a critical component of the voter's identity.
USVoterData_BF__data__Pennsylvania__2015__blackhole_ak_orig__records.csv13 columns487,414 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 8 | firstName | high | [8] header 'fname', values are common given names (Jennifer, Charlotte, Susheela, Carl, Abbey) |
| 9 | middleName | high | [9] header 'mname', values are single-letter middle initials (M, A, L) |
| 10 | lastName | high | [10] header 'lname', values are surnames (Kain, Genovese, Roach, Christiansen, Russell) |
| 11 | suffix | high | [11] header 'suffix', value 'Jr' is a generational suffix |
| 12 | gender | high | [12] header 'sex', values are M/F gender codes |
| 20 | phone | high | [20] header 'phone', values are 10-digit phone numbers (9072245244, 9072245383) |
| 23 | address1 | high | [23] header 'address', values are street addresses (12519 Merlin Dr, 32320 Blying Sound Dr Harborview) |
| 25 | city | high | [25] header 'city', values are city names (Seward, Soldotna) |
| 28 | zip | high | [28] header 'zip', values are 5-digit zip codes (99664, 99669) |
| 30 | address1 | high | [30] header 'mail_address', values are mailing street addresses (Po Box 1122, Po Box 2233, 2493 Fm 949) |
| 32 | city | high | [32] header 'mail_city', values are mailing city names (Seward, Cat Spring) |
| 33 | state | high | [33] header 'mail_state', values are US state codes (AK, TX) |
| 34 | zip | high | [34] header 'mail_zip', values are mailing zip codes (99664, 78933) |
Notes: US voter registration data with 93 columns total; 13 contain PII. Column [5] 'state' contains voter registration state (AK) but is not residence state — not mapped as state field since it's administrative. Column [13] 'dob' contains only '0000-00-00' (nulled/redacted values), not valid DOB data. Column [19] 'email' is empty across all samples. Columns 36-48 and 50-88 are administrative/district codes and voting history flags — skipped as non-PII.
USVoterData_BF__data__Pennsylvania__2015__blackhole_al__records.csv18 columns132,787 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 5 | state | high | [5] header 'state', values are US state abbreviations (AL) |
| 8 | firstName | high | [8] header 'fname', values are common given names (Rose, Laura, Kristie) |
| 9 | middleName | high | [9] header 'mname', values are middle names/initials (Marie, B, Lynn Wallace) |
| 10 | lastName | high | [10] header 'lname', values are common surnames (Allen, Aldridge, Alexander) |
| 11 | suffix | high | [11] header 'suffix', generational/credential suffix field after name |
| 12 | gender | high | [12] header 'sex', values are M/F gender codes |
| 13 | dob | high | [13] header 'dob', values are YYYY-MM-DD birth dates (1961-02-04, 1950-04-05) |
| 19 | skip | high | [19] header 'email', email address field |
| 20 | phone | high | [20] header 'phone', values are 10-digit phone numbers (2562333566) |
| 23 | address1 | high | [23] header 'address', values are street addresses (14101 Seven Mile Post Rd) |
| 24 | address2 | high | [24] header 'address2', secondary address line field |
| 25 | city | high | [25] header 'city', values are city names (Athens, Elkmont) |
| 28 | zip | high | [28] header 'zip', values are 5-digit US ZIP codes (35611, 35620) |
| 30 | address1 | high | [30] header 'mail_address', values are mailing street addresses (14101 SEVEN MILE POST Rd) |
| 31 | address2 | high | [31] header 'mail_address2', secondary mailing address line field |
| 32 | city | high | [32] header 'mail_city', values are mailing city names (Athens, Elkmont) |
| 33 | state | high | [33] header 'mail_state', values are US state abbreviations (AL) |
| 34 | zip | high | [34] header 'mail_zip', values are 5-digit US ZIP codes (35611, 35620) |
Notes: 93 columns total; voter registration compilation covering AL and other states. Columns 0-3 are internal/system IDs (skip). Column 7 'prefix' appears to be honorific salutation (skip). Columns 50-88 are election participation flags by year (skip). Voting history, precinct codes, district assignments, fips, and source_file are all non-PII administrative fields (skip). Column 27 'residential_state' has no sample values but maps conceptually to state — skipped due to no values confirming PII content. Mail address fields (30-34) mapped as additional address instances alongside residential address fields (23-25, 28).
USVoterData_BF__data__Pennsylvania__2018__Info__Political_Party_Codes_and_Descriptions.txt0 rows
File structure
Notes: The provided data is a free-form text listing political party codes and descriptions, not a structured dataset with columns. It contains no discernible PII fields or column structure, and appears to be a reference table or legend for party affiliations rather than actual voter records.
USVoterData_BF__data__Pennsylvania__2018__Statewide__ADAMS__ADAMS_Election_Map_20181001.txt0 rows
File structure
Notes: free-form text with no consistent delimiter or row structure — appears to be a list of election participation records for a single individual ("ADAMS") with election type and date, not a CSV with PII columns. No usable structured columns present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__ADAMS__ADAMS_FVE_20181001.txt14 columns66,486 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header-less but values are surnames (LANTZ, WEIKERT, STOUFFER, ADKINS, SMITH, RISS) |
| 3 | firstName | high | [3] header-less but values are common given names (RICHARD, MARK, JOHN, ELLEN, JACK, JEAN) |
| 4 | middleName | medium | [4] header-less single letters (G, K, C, E, A, A) typical for middle initials |
| 5 | suffix | medium | [5] values like SR (senior generational suffix) |
| 6 | gender | high | [6] values M/F clearly indicating gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, represents birth date (05/16/1923, 11/20/1932, etc.) |
| 8 | skip | high | [8] additional DOB column with year-only dates (01/01/1955, 01/01/1959, etc.) |
| 10 | skip | high | [10] registration date or additional DOB component (01/11/2017, 09/06/2017, etc.) |
| 14 | address1 | high | [14] values are street addresses (HALLECK DR, FULTON DR, CHAMBERSBURG RD, CARLISLE PIKE, COON RD) |
| 15 | address2 | low | [15] apartment/unit numbers where present (#105) |
| 17 | city | high | [17] values are city names (EAST BERLIN, NEW OXFORD, FAYETTEVILLE, ASPERS, BIGLERVILLE) |
| 18 | state | high | [18] consistent state abbreviation (PA) |
| 19 | zip | high | [19] 5-digit zip codes (17316, 17350, 17222, 17350, 17304, 17307) |
| 150 | skip | high | [150] 10-digit phone numbers (7173522009, 7173371519) |
Notes: 153 total columns, 12 contain PII. Columns 0,1,13,20-152 contain internal IDs, flags, timestamps, or empty values and are skipped. This is a US voter registration dataset with detailed personal information including full names, addresses, birth dates, gender, and phone numbers.
USVoterData_BF__data__Pennsylvania__2018__Statewide__ADAMS__ADAMS_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains only geographic and administrative codes (county, district, township, etc.) with no personal identifiable information. All columns represent location identifiers and district numbers, not PII fields. No emails, phone numbers, names, addresses, DOB, SSN, or other PII are present in the sample data.
USVoterData_BF__data__Pennsylvania__2018__Statewide__ADAMS__ADAMS_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header-less metadata file describing geographic and administrative divisions (precincts, wards, districts, municipalities, etc.) rather than containing personal voter records. The sample shows repeated entries for 'ADAMS' with codes and descriptions of political/geographic entities. No PII fields are present in this structure — this appears to be a reference table for voter data files, not the actual voter data itself. Therefore, no columns map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__ALLEGHENY__ALLEGHENY_Election_Map_20181001.txt0 rows
File structure
Notes: free-form text with election metadata, no structured columns or PII visible in the first 50 rows
USVoterData_BF__data__Pennsylvania__2018__Statewide__ALLEGHENY__ALLEGHENY_FVE_20181001.txt12 columns931,944 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] values are surnames (BAUM, BERKHEIMER, BROUJOS, etc.) |
| 3 | firstName | high | [3] values are first names (GRACE, MARTHA, LOUISE, PAUL, etc.) |
| 4 | middleName | high | [4] single initial values (W, E, J, L, BRADLEY) represent middle names |
| 6 | gender | high | [6] values are gender codes (F, M) |
| 7 | skip | high | [7] values match MM/DD/YYYY birth date format |
| 8 | skip | high | [8] additional birth date column in MM/DD/YYYY format |
| 14 | address1 | high | [14] contains street address values (FOX CHAPEL RD, MERIDIAN RD, PORTLAND ST, etc.) |
| 15 | address2 | medium | [15] apartment/unit numbers where present (116A, 502, 2301, etc.) |
| 17 | city | high | [17] contains city names (PITTSBURGH, GIBSONIA, SEWICKLEY, etc.) |
| 18 | state | high | [18] all values are 'PA' (Pennsylvania) |
| 19 | zip | high | [19] contains 5-digit ZIP codes (15238, 15044, 15206, etc.) |
| 150 | skip | medium | [150] contains 10-digit phone numbers (4123628025, 7178776312) |
Notes: 153 total columns identified; 12 contain PII. Columns 0,1,5,9-13,16,20-149 (except 150) are internal IDs, party codes, district codes, or empty fields and should be skipped per rules. This is a standard US voter registration format with detailed personal information.
USVoterData_BF__data__Pennsylvania__2018__Statewide__ALLEGHENY__ALLEGHENY_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a structured file with consistent tab-delimited format, but all visible columns represent geographic identifiers (county, precinct codes, and place names). No PII fields such as names, addresses, dates of birth, or contact information are present in the sample. The data appears to be purely geographic and administrative boundaries, likely used for voting district mapping. No emails, phone numbers, or personal identifiers are visible.
USVoterData_BF__data__Pennsylvania__2018__Statewide__ALLEGHENY__ALLEGHENY_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: yes·Quote: "
Notes: This file appears to be a header-only or metadata file describing voting districts and codes. No PII fields are present in the first 50 rows. All columns contain geographic/district codes and descriptions, not personal voter data. This is likely a reference file for interpreting voter registration data, not the actual voter records themselves.
USVoterData_BF__data__Pennsylvania__2018__Statewide__ARMSTRONG__ARMSTRONG_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a structured file with no PII columns present. The data appears to be election event records (election names, IDs, and dates) with no personal identifying information. All columns are non-PII according to exclusion rules (election metadata, IDs, and dates).
USVoterData_BF__data__Pennsylvania__2018__Statewide__ARMSTRONG__ARMSTRONG_FVE_20181001.txt14 columns41,088 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header suggests last name, values are surnames like FORESTER, RUPP, WATT |
| 3 | firstName | high | [3] header suggests first name, values are given names like SANDRA, ROY, TIMOTHY |
| 4 | middleName | high | [4] header suggests middle name or initial, values are single letters F, J, A, K |
| 5 | suffix | high | [5] values include JR, mapping to suffix field |
| 6 | gender | high | [6] values are F/M/U mapping to gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern and represent birth dates |
| 8 | skip | high | [8] additional date column likely alternate birth date format |
| 14 | address1 | high | [14] values are street addresses like MCHADDON RD, STATE ROUTE 1042 |
| 15 | address2 | medium | [15] contains apartment/unit numbers like 8, 132 |
| 16 | address2 | medium | [16] contains PO boxes like P.O.BOX 93, PO BOX 88 |
| 17 | city | high | [17] values are city names like KITTANNING, NUMINE, FORD CITY |
| 18 | state | high | [18] values are state abbreviations like PA |
| 19 | zip | high | [19] values are 5-digit zip codes like 16201, 16244 |
| 150 | skip | high | [150] contains 10-digit phone numbers like 7247835053 |
Notes: 153 total columns identified; 12 contain PII (names, addresses, DOB, gender, phone). Remaining columns appear to be internal voter IDs, precinct codes, party affiliations, and other administrative data which are skipped per rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__ARMSTRONG__ARMSTRONG_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The file contains purely geographic/political administrative divisions (boroughs, townships, districts) with no personal identifiable information. All columns represent location codes and names rather than individual voter data. No PII fields are present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__ARMSTRONG__ARMSTRONG_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header description file, not actual voter records. It contains only metadata about column meanings (e.g., 'Precinct', 'School district'). No PII fields are present in these rows. The actual data would contain voter registration details, but this sample only shows the schema legend.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BEAVER__BEAVER_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a non-PII structured file listing election events with consistent tab-delimited format. Columns contain state codes, event IDs, election descriptions, and dates — no personal identifiers present. All rows follow the same pattern with no embedded emails, phone numbers, or other PII.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BEAVER__BEAVER_FVE_20181001.txt10 columns109,695 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header implies last name, values are common surnames |
| 3 | firstName | high | [3] header implies first name, values are common given names |
| 4 | middleName | medium | [4] values are single initials/letters, consistent with middle name representation |
| 6 | gender | high | [6] values are 'M'/'F' which map directly to gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern and represent birth dates |
| 14 | address1 | high | [14] header indicates address, values are street addresses |
| 17 | city | high | [17] header and values clearly represent city names |
| 18 | state | high | [18] values are standard US state abbreviations (PA) |
| 19 | zip | high | [19] values are 5-digit zip codes |
| 150 | skip | high | [150] values are 10-digit numbers matching US phone number format |
Notes: 153 columns total, 10 contain PII: lastName, firstName, middleName, gender, dob, address1, city, state, zip, phone. Remaining columns are internal IDs, geographic codes, voter status flags, and timestamps which are skipped per rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BEAVER__BEAVER_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The file is not a delimited dataset with PII columns; it's a structured listing of Pennsylvania voter precincts for Beaver County with no personal information. Each row contains a county name, a numeric code, and a precinct name/township, none of which constitute PII. No emails, addresses, names, or other sensitive data are present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BEAVER__BEAVER_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data is a header-less metadata row describing geographic/political district codes for Beaver County voter records. No PII fields appear in this row; all entries are codes, labels, and placeholder values (NU = Not Used). Subsequent rows likely contain actual voter records with PII. This sample only includes the descriptive header row, so no PII columns can be mapped from it alone.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BEDFORD__BEDFORD_Election_Map_20181001.txt0 rows
File structure
Notes: free-form text with no consistent delimiter or structured columns; appears to be a list of election events with location and dates rather than personal voter records
USVoterData_BF__data__Pennsylvania__2018__Statewide__BEDFORD__BEDFORD_FVE_20181001.txt12 columns31,260 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header indicates last name, values are surnames like STOEFFLER, GATES |
| 3 | firstName | high | [3] header indicates first name, values are common given names: JOHN, SUSAN, TIMOTHY |
| 4 | middleName | high | [4] header and values align with middle name conventions (single letters M, A, D, W) |
| 6 | gender | high | [6] contains gender codes M/F matching voter data patterns |
| 7 | skip | high | [7] format MM/DD/YYYY matches birth date patterns in voter records |
| 8 | skip | high | [8] second birth date column with MM/DD/YYYY format |
| 14 | address1 | high | [14] contains street address components like FRIENDSHIP VILLAGE RD |
| 17 | city | high | [17] contains city names: BEDFORD, SCHELLSBURG, JAMES CREEK |
| 18 | state | high | [18] contains consistent state abbreviation PA |
| 19 | zip | high | [19] contains valid 5-digit ZIP codes: 15522, 15559 |
| 150 | skip | high | [150] contains 10-digit US phone numbers: 7175762395, 8146352438 |
| 151 | city | high | [151] duplicate city information matching column 17 |
Notes: 153 total columns identified; 11 contain PII including duplicate city field. All party affiliation/political codes (columns 9,11,73,75,etc) are skipped per exclusion rules for internal flags
USVoterData_BF__data__Pennsylvania__2018__Statewide__BEDFORD__BEDFORD_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The file appears to be a structured list of geographic districts and regions within Bedford County, Pennsylvania. It contains only geographic identifiers, district codes, and region descriptions with no personal identifying information (PII). All columns represent location codes and area names rather than individual voter data.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BEDFORD__BEDFORD_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data is purely structural metadata describing geographic and administrative divisions (Precinct, City ward, School district, Municipality, Magistrate, Legislative, Senate, Congressional, School district region, County). No personal identifying information (PII) fields like names, addresses, dates of birth, or contact details appear in the visible rows. All columns represent non-PII geographic codes and labels, thus no PII columns to map.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BERKS__BERKS_Election_Map_20181001.txt0 rows
File structure
Notes: free-form text with no consistent delimiter or structured columns; contains election cycle records and dates, not PII
USVoterData_BF__data__Pennsylvania__2018__Statewide__BERKS__BERKS_FVE_20181001.txt10 columns254,171 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header implies last name, values are common surnames |
| 3 | firstName | high | [3] header implies first name, values are common given names |
| 4 | middleName | high | [4] single initial values, column name suggests middle name component |
| 6 | gender | high | [6] values are 'M'/'F'/'U', header implies gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, header implies birth date |
| 14 | address1 | high | [14] values are street addresses, header implies address line 1 |
| 17 | city | high | [17] values are city names, header implies city |
| 18 | state | high | [18] values are two-letter state abbreviations, header implies state |
| 19 | zip | high | [19] values are 5-digit zip codes, header implies zip/postal code |
| 150 | skip | high | [150] values are 10-digit numbers, header implies phone number |
Notes: 153 columns total, 10 contain PII (names, addresses, DOB, gender, phone). Remaining columns are internal IDs, geographic codes, party affiliation, voting status, and other non-PII administrative fields. All PII columns mapped with high confidence based on header names and sample values.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BERKS__BERKS_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The file contains purely geographic codes and location names (precinct identifiers). No personal identifiers such as names, addresses, or contact information are present. All entries are structured as county identifiers, codes, and location names, which do not qualify as PII under the defined fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BERKS__BERKS_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a metadata header-only file describing voter registration fields for Berks County. No actual voter records with PII are present in the first 50 rows. All visible columns represent geographic/district codes, administrative divisions, and structural fields — none contain personal identifying information. The data appears to be a schema or field definition listing rather than actual voter records.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BLAIR__BLAIR_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains unstructured election participation records with no personal identifying information. Each line represents a voter named BLAIR participating in a specific election with a date. There are no columns for names, addresses, DOB, phone numbers, etc. — only static voter name "BLAIR" and election metadata. No PII fields present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BLAIR__BLAIR_FVE_20181001.txt12 columns75,534 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header implies last name, values are common surnames |
| 3 | firstName | high | [3] header implies first name, values are common given names |
| 4 | middleName | medium | [4] values include both middle names and honorifics; mapped as middleName per values like SHELLITO, JOHN, JOHN |
| 6 | gender | high | [6] values are 'F'/'M' which map to gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern and represent birth dates |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern and represent birth dates |
| 14 | address1 | high | [14] header and values clearly represent street addresses |
| 16 | city | high | [16] header and values represent city names |
| 17 | city | high | [17] values are city names like MARTINSBURG, DUNCANSVILLE, ALTOONA |
| 18 | state | high | [18] values are two-letter US state abbreviations like PA |
| 19 | zip | high | [19] values are 5-digit US ZIP codes |
| 150 | skip | high | [150] values are 10-digit US phone numbers |
Notes: 153 total columns; 9 contain PII. Columns 4 contains mixed honorifics and middle names; mapped as middleName based on majority values. Columns with party affiliation codes (e.g., D/R) and precinct codes were excluded per rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BLAIR__BLAIR_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data consists of structured voting district codes and identifiers rather than personal PII fields. All visible columns contain geographic identifiers, district codes, and administrative divisions without any names, addresses, contact details, or other PII. This is typical voter roll metadata used for precinct mapping, not individual voter records. No PII fields are present to map.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BLAIR__BLAIR_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header row for a US voter data file. The first row contains placeholder values ('BLAIR') and descriptive labels for each column. There are no PII values in this row. Subsequent rows likely contain actual voter data with PII fields such as voter_id, first_name, last_name, address, city, state, zip, dob, gender, etc. Further analysis of rows 14+ is required to map PII columns accurately.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BRADFORD__BRADFORD_Election_Map_20181001.txt0 rows
File structure
Notes: free-form text with no consistent column structure - appears to be a list of election events with election type and date, no personal identifiers present in the sample
USVoterData_BF__data__Pennsylvania__2018__Statewide__BRADFORD__BRADFORD_FVE_20181001.txt11 columns36,210 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header suggests last name, values are surnames like SHOLLY, BONNERWITH, GELBUTIS |
| 3 | firstName | high | [3] header suggests first name, values are common given names like VICTOR, NANCY, GEORGIA |
| 4 | middleName | high | [4] header suggests middle name or initial, values are single letters like N, ANN, C, LEE |
| 6 | gender | high | [6] values are M/F, typical gender codes |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, header suggests birth date |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern, another date field likely secondary birth date reference |
| 14 | address1 | high | [14] values are street addresses like FRANCIS ST, SHESHEQUIN RD, STOVAK RD |
| 17 | city | high | [17] values are city names like SAYRE, ULSTER, ATHENS, TROY |
| 18 | state | high | [18] values are state abbreviations like PA |
| 19 | zip | high | [19] values are 5-digit zip codes like 18840, 18850, 18810 |
| 150 | skip | low | [150] contains a 10-digit number 5702055932 which matches US phone number format |
Notes: 153 total columns identified; 11 contain PII (names, gender, DOB, address components, phone). Remaining columns are internal IDs, district/county codes, registration metadata, and empty fields — all skipped per rules. Data matches US voter registration structure with detailed personal and geographic information.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BRADFORD__BRADFORD_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains purely geographic and administrative codes without any personal identifying information. It lists municipalities, school districts, judicial districts, and other jurisdictional codes for Bradford County, PA. No names, addresses, or PII fields are present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BRADFORD__BRADFORD_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: yes·Quote: "
Notes: This is a header/metadata row structure, not actual voter data. The first row appears to be a header label row ("BRADFORD"), and subsequent rows define codes and descriptions for voting districts/precincts (WARD, SC, MN, etc.). No PII is present in these rows — they are purely administrative codes and descriptions. The actual voter records (names, addresses, DOB, etc.) would appear later in the file and need to be processed separately.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BUCKS__BUCKS_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file appears to be a structured table of election dates and identifiers, but contains no personal identifiable information (PII). All columns represent election-related metadata (county codes, election types, and dates) which are not classified as PII per the rules. Therefore, no columns map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BUCKS__BUCKS_FVE_20181001.txt13 columns452,482 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header indicates last name, values are common surnames |
| 3 | firstName | high | [3] header indicates first name, values are common given names |
| 4 | suffix | high | [4] values are generational suffixes (J, D, W, E, A) |
| 6 | gender | high | [6] values are 'M'/'F' which map to gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern and represent birth dates |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern and represent birth dates |
| 14 | address1 | high | [14] header indicates address line, values are street addresses |
| 15 | address2 | medium | [15] values appear to be apartment/unit numbers or additional address information |
| 17 | city | high | [17] header indicates city, values are city names |
| 18 | state | high | [18] header indicates state, values are state abbreviations (PA) |
| 19 | zip | high | [19] header indicates zip, values are 5-digit zip codes |
| 92 | partyAffiliation | — | [92] values represent political party affiliation (AP, D, R, NOP), not PII |
| 151 | county | — | [151] values represent county names (BUCKS), not PII |
Notes: 153 columns total, 9 contain PII (lastName, firstName, suffix, gender, dob (x2), address1, address2, city, state, zip). Additional columns represent political party affiliation and county which are not considered PII per exclusion rules. The data follows a structured CSV format with consistent columns across all rows.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BUCKS__BUCKS_Zone_Codes_20181001.txt0 rows
File structure
Notes: free-form text with embedded precinct codes, no structured columns
USVoterData_BF__data__Pennsylvania__2018__Statewide__BUCKS__BUCKS_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a headerless metadata file describing voting districts and codes, not actual voter records. Columns contain geographic codes (Precinct, School district, Municipal, etc.) and administrative numbers — no personal identifying information present. All columns map to skip per exclusion rules (internal codes, geographic identifiers, not PII).
USVoterData_BF__data__Pennsylvania__2018__Statewide__BUTLER__BUTLER_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII columns detected. File appears to contain election history records with voter ID 'BUTLER' and election details (election type and date). No personal identifiable information such as names, addresses, phone numbers, or dates of birth is present in the first 50 rows.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BUTLER__BUTLER_FVE_20181001.txt14 columns127,353 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header implies last name, values are surnames like 'BITTING', 'URLING', 'TORSELL' |
| 3 | firstName | high | [3] header implies first name, values are given names like 'ROMAINE', 'ELIZABETH', 'CAROL' |
| 4 | middleName | high | [4] header implies middle name, values are initials or names like 'P', 'JAYNE', 'S' |
| 5 | suffix | high | [5] values are honorific suffixes like 'SR' |
| 6 | gender | high | [6] values are 'F'/'M' mapping to gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, header suggests birth date |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern, likely alternate birth date field |
| 14 | address1 | high | [14] values are street addresses like 'MARWOOD RD', 'JAMESTOWN CT' |
| 15 | zip | high | [15] values are 5-digit ZIP codes |
| 16 | city | high | [16] values are city names like 'CABOT', 'VALENCIA', 'BUTLER' |
| 17 | state | high | [17] values are state abbreviations like 'PA' |
| 18 | state | high | [18] duplicate state field |
| 19 | zip | high | [19] values are 5-digit ZIP codes |
| 150 | skip | high | [150] values are 10-digit phone numbers |
Notes: 153 columns total, 13 contain PII. Columns 0,1,13,20-22,24-32,34-42,44-49,51-59,61-69,71-73,75-85,87-89,91-93,95-99,101-105,107-109,111-113,115-119,121-125,127-129,131-133,135-137,139-149,151-152 are internal IDs, timestamps, geographic codes, or political affiliations and were skipped.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BUTLER__BUTLER_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a structured file containing geographic and administrative codes (precincts, school districts, legislative districts) rather than personal voter records. No PII fields (names, addresses, emails, etc.) are present in the sample data. The file appears to be a reference table of district codes for Butler County, PA, used for organizing voter data.
USVoterData_BF__data__Pennsylvania__2018__Statewide__BUTLER__BUTLER_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file appears to be a headerless lookup table for geographic codes (Precinct, Ward, School District, etc.) and contains no personal identifiable information. All columns represent internal codes and geographic designations, not PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CAMBRIA__CAMBRIA_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains election event metadata (county name, event number, event description, and date) with no personal identifiable information. All columns represent event identifiers and dates rather than individual voter data. No PII fields are present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CAMBRIA__CAMBRIA_FVE_20181001.txt6 columns0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header 'NAGLE', values are common surnames |
| 3 | firstName | high | [3] header 'MARY', values are common given names |
| 4 | middleName | medium | [4] values like 'THERESA' and initials 'L', 'M', 'J', 'A' indicate middle name or maiden name |
| 6 | gender | high | [6] values are 'F' and 'M' — unambiguous gender codes |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern and represent birth dates |
| 150 | skip | medium | [150] values are 10-digit US phone numbers |
Notes: 153 columns total; only columns 2, 3, 4, 6, 7, and 150 contain PII. Columns 0, 1, 5-149 are internal IDs, empty fields, or political/registration codes and are skipped per rules. No company/business fields present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CAMBRIA__CAMBRIA_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is not structured data with PII columns. The file appears to be a list of geographic codes and place names (townships, boroughs, schools) for Cambria County, Pennsylvania voter registration data. There are no personal identifiers (names, addresses, emails, etc.) in the visible rows. The format is tab-delimited with no header row. This is likely part of a voter registration dataset where the actual PII would be in other columns not shown in this sample.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CAMBRIA__CAMBRIA_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header-only row showing column mappings for a voter data file. No actual voter records are present in the first 50 rows — only the column name mappings. All columns are administrative codes (Precinct, Ward, School district, etc.) and contain no PII. Therefore, no PII columns to map.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CAMERON__CAMERON_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains only election-related data (voter registration history) with no PII fields present. Each row shows a voter ID (CAMERON) and election details (election type and date). No personal identifying information such as names, addresses, DOB, etc., is included in the first 50 rows provided.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CAMERON__CAMERON_FVE_20181001.txt10 columns2,907 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header 'lastName' inferred from values like 'WIRT', 'GUISE', 'KANE' |
| 3 | firstName | high | [3] header 'firstName' inferred from values like 'DAWN', 'REBECCA', 'ABIGAIL' |
| 4 | middleName | medium | [4] values include single letters and full names; header suggests middle name |
| 6 | gender | high | [6] values are 'F'/'M' which map directly to gender |
| 7 | skip | high | [7] values follow MM/DD/YYYY date format, typical for birth dates |
| 14 | address1 | high | [14] contains street address components like 'JERICHO RD', 'STERLING RUN RD' |
| 17 | city | high | [17] values are city names like 'SINNEMAHONING', 'DRIFTWOOD', 'EMPORIUM' |
| 18 | state | high | [18] all values are 'PA' indicating Pennsylvania |
| 19 | zip | high | [19] contains 5-digit ZIP codes like '15861', '15832', '15834' |
| 150 | skip | high | [150] contains 10-digit phone numbers like '8146010758', '8145462368' |
Notes: 153 total columns; 10 contain PII (names, DOB, gender, address components, phone). Remaining columns are internal IDs, geographic codes, party affiliation, and timestamps — all excluded per rules. Data matches US voter registration structure with state PA focus.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CAMERON__CAMERON_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is unstructured voter registration data showing only geographic identifiers (ward, township, district codes) and county names. No PII fields (names, addresses, DOB, etc.) appear in the first 50 rows. The format is tab-delimited but contains only administrative codes, not personal voter records. Therefore, no PII columns to map.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CAMERON__CAMERON_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header-less metadata file describing geographic districts and precincts. No PII fields are present. All rows appear to be structural descriptors (e.g., 'Precinct', 'School district', 'Municipality'). The repeated 'CAMERON' value likely denotes a location label rather than personal data. No columns map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CARBON__CARBON_Election_Map_20181001.txt0 rows
File structure
Notes: free-form text with no consistent delimiter structure; appears to be a list of election events rather than structured voter records
USVoterData_BF__data__Pennsylvania__2018__Statewide__CARBON__CARBON_FVE_20181001.txt12 columns42,873 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header suggests last name, values are surnames |
| 3 | firstName | high | [3] header suggests first name, values are given names |
| 4 | middleName | medium | [4] sparse but when present values appear as middle names or initials |
| 5 | suffix | low | [5] single value 'III' matches generational suffix pattern |
| 6 | gender | high | [6] values 'M'/'F' clearly indicate gender |
| 7 | skip | high | [7] values match MM/DD/YYYY birth date format |
| 8 | skip | high | [8] additional birth date column in same format |
| 14 | address1 | high | [14] values are street addresses |
| 15 | zip | high | [15] 5-digit postal codes |
| 17 | city | high | [17] values are city names |
| 18 | state | high | [18] consistent 'PA' state abbreviation |
| 150 | skip | high | [150] 10-digit US phone numbers |
Notes: 153 total columns; 12 contain PII (names, DOB, gender, address, phone). Remaining columns are internal IDs, voting districts, party affiliations, status flags, and other non-PII administrative data. File follows standard US voter registration layout with multiple date columns (registration, election participation) and geographic district codes.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CARBON__CARBON_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data is purely geographic and administrative codes (county, district, township, school district) with no personal identifying information present. All entries are location identifiers and codes, not PII.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CARBON__CARBON_Zone_Types_20181001.txt0 rows
File structure
Notes: This is not PII data - it's a header mapping file showing column purposes for a voter data CSV. The rows contain only codes and descriptions (e.g., CARBON, Precinct, School district). No actual voter records or PII values are present in these first 50 rows. The real data would appear later in the file with actual voter names, addresses, etc.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CENTRE__CENTRE_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data appears to be a list of election events with no personal identifiable information (PII). The columns contain election names and dates, which do not map to any PII fields. Therefore, no columns are mapped to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CENTRE__CENTRE_FVE_20181001.txt17 columns109,426 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header suggests surname, values are common last names: 'ERNST', 'MILLER', 'KLEMICK', 'JOHNSEN', 'MOYER' |
| 3 | firstName | high | [3] header suggests given name, values are common first names: 'JAYNE', 'SUSAN', 'REID', 'MARY', 'AMANDA', 'ROBERT' |
| 4 | middleName | medium | [4] values include both single letters and full names; maps to middleName as per pattern |
| 6 | gender | high | [6] values are 'F'/'M' which map directly to gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern and represent birth dates: '12/31/1925', '01/18/1942', etc. |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern and represent birth dates: '01/01/1980', '01/01/1966', etc. |
| 10 | skip | high | [10] values match MM/DD/YYYY date pattern and represent registration or enrollment dates: '01/11/2016', '10/05/2005', etc. |
| 14 | address1 | high | [14] values are street addresses: 'BLUE SPRING LN', 'HONORS LN', 'LIONS HILL RD', etc. |
| 17 | city | high | [17] values are city names: 'BOALSBURG', 'STATE COLLEGE', 'LEMONT' |
| 18 | state | high | [18] values are two-letter state abbreviations: 'PA' |
| 19 | zip | high | [19] values are 5-digit zip codes: '16827', '16803', '16851' |
| 20 | address2 | medium | [20] values include 'P.O. BOX' which maps to address line 2 |
| 22 | city | high | [22] duplicate city values: 'BOALSBURG', 'LEMONT' |
| 23 | state | high | [23] duplicate state values: 'PA' |
| 24 | zip | high | [24] duplicate zip codes: '16827', '16851' |
| 25 | skip | high | [25] values match MM/DD/YYYY date pattern and represent registration dates: '11/04/2008', '05/15/2018', etc. |
| 150 | skip | high | [150] values are 10-digit numbers: '8142388120', '7173859164' — typical US phone format |
Notes: 153 total columns; 21 contain PII (names, dates of birth, addresses, phone). Remaining columns are internal voter IDs, precinct codes, party affiliations, status flags, and other administrative data — all skipped per rules. Multiple date columns represent different voter lifecycle events (birth, registration, last vote, etc.) and all map to dob per pattern.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CENTRE__CENTRE_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The file contains only geographic and administrative codes (municipalities, townships, school districts, legislative districts). No personal identifiable information (PII) fields such as names, addresses, or contact details are present. All columns represent location identifiers and district codes, which do not qualify as PII per the provided definitions.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CENTRE__CENTRE_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file is a metadata header describing voting districts and regions, not actual voter records. It contains only geographic and administrative codes (CENTRE, Precinct, Municipality, School district, etc.) with no personal identifiers. No PII fields present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CHESTER__CHESTER_Election_Map_20181001.txt0 rows
File structure
Notes: free-form text with no consistent delimiter or column structure; contains election participation data with no PII fields
USVoterData_BF__data__Pennsylvania__2018__Statewide__CHESTER__CHESTER_FVE_20181001.txt11 columns353,140 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] values are common surnames: 'DIETRICK', 'MABIUS', 'WISS', 'GREEN', 'JONES' |
| 3 | firstName | high | [3] values are common given names: 'HELEN', 'GERALDINE', 'RAYMOND', 'JOSEPH', 'TERRY', 'SALLY' |
| 4 | middleName | medium | [4] contains initials and some full middle names: 'W', 'J', 'E', 'L', 'ANNE', 'C' |
| 6 | gender | high | [6] values are gender codes: 'F', 'M' |
| 7 | skip | high | [7] values match MM/DD/YYYY date format, interpreted as birth dates |
| 8 | skip | high | [8] additional birth date column in MM/DD/YYYY format |
| 14 | address1 | high | [14] contains street addresses: 'STANTON AVE', 'N VALLEY FORGE RD', 'FREEDOM BLVD', etc. |
| 17 | city | high | [17] contains city names: 'WEST CHESTER', 'DEVON', 'COATESVILLE', etc. |
| 18 | state | high | [18] consistent state abbreviation 'PA' (Pennsylvania) |
| 19 | zip | high | [19] 5-digit postal codes: '19382', '19333', '19320', etc. |
| 150 | skip | high | [150] 10-digit numbers interpreted as phone numbers: '6106420576', '7177614896', '4846678625' |
Notes: 153 columns total, 11 contain PII. This is a US voter registration dataset with detailed personal information including names, addresses, birth dates, gender, and phone numbers. Columns 0, 1, 5, 9-13, 15-23, 25-152 appear to be internal IDs, precinct codes, status flags, or empty fields and are mapped as skip.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CHESTER__CHESTER_Zone_Codes_20181001.txt0 rows
File structure
Notes: This file is purely geographic data with no PII fields. It contains voting precinct identifiers (names and codes) but no personal information like names, addresses, or contact details. No emails or phone numbers are visible in the sample.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CHESTER__CHESTER_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is not structured PII data; it's a header-less reference table defining geographic codes and district abbreviations used in voter rolls. All rows describe administrative divisions (Precinct, Ward, School district, etc.) and contain no personal identifiers. No columns map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CLARION__CLARION_Election_Map_20181001.txt0 rows
File structure
Notes: This file is not structured data; it's a free-form text listing election events with no consistent column pattern. It contains no PII fields like names, addresses, or identifiers. The format appears to be a simple list of election types and dates rather than a delimited dataset with personal information.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CLARION__CLARION_FVE_20181001.txt12 columns22,717 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header 'lastName' (implied by position and values), values are common surnames |
| 3 | firstName | high | [3] header 'firstName' (implied by position and values), values are common given names |
| 4 | middleName | high | [4] header 'middleName' (implied by position and values), values are single letters or short names |
| 6 | gender | high | [6] values are 'M'/'F'/'U', header implies gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, header implies birth date |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern, likely second DOB column (possibly registration date but format matches DOB) |
| 14 | address1 | high | [14] values are street addresses (e.g., 'YEANY LN', 'DOMENICA CIR') |
| 17 | city | high | [17] values are city names (e.g., 'MAYPORT', 'CLARION') |
| 18 | state | high | [18] values are state abbreviations (e.g., 'PA') |
| 19 | zip | high | [19] values are 5-digit zip codes |
| 150 | skip | high | [150] values are 10-digit numbers, likely phone numbers |
| 151 | city | high | [151] duplicate city values, same as column 17 |
Notes: 153 columns total, 12 contain PII. Columns 0, 1, 5, 9, 10, 11-152 contain non-PII data such as internal IDs, party affiliation, precinct codes, and other administrative/voting records. Column 150 contains phone numbers despite no explicit header.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CLARION__CLARION_Zone_Codes_20181001.txt0 rows
File structure
Notes: free-form text with embedded location codes, no consistent columnar structure for PII fields
USVoterData_BF__data__Pennsylvania__2018__Statewide__CLARION__CLARION_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: yes·Quote: "
Notes: This appears to be a metadata header file describing voter registration fields, not actual voter records. The sample rows contain only field descriptions (e.g., 'Precinct', 'City ward') and no personal data. No PII fields are present in this snippet.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CLEARFIELD__CLEARFIELD_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file appears to be a structured table of election events with no personal identifying information. Columns contain election names, sequence numbers, and election dates — all non-PII. No names, addresses, voter IDs, or other PII appear in the sample. Therefore, no columns map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CLEARFIELD__CLEARFIELD_FVE_20181001.txt14 columns46,626 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] values match common surnames (BRIGHTBILL, FETTERHOFF, KNEPP, DANNER, REDDEN) |
| 3 | firstName | high | [3] values match common given names (GENEVIEVE, SONYA, STEVEN, ROBERT, HERBERT, MANDY) |
| 4 | middleName | medium | [4] contains both blank and single-letter values (SUE, MARK, RUSSELL, M) which can be middle initials or names |
| 5 | suffix | medium | [5] contains generational suffix 'JR' |
| 6 | gender | high | [6] values are gender codes (F, M) |
| 7 | skip | high | [7] date format MM/DD/YYYY with plausible birth years (1954-1958) |
| 14 | address1 | high | [14] contains street addresses (PINDER POINT RD, GRANDE TERRE CT, DOUGLAS RD, MANN RD, FRIENDLY ACRES RD) |
| 15 | city | medium | [15] contains city name 'LAKE' |
| 16 | address2 | medium | [16] contains secondary address information (1931 treasure lake, 137 TREASURE LAKE, P O BOX 81) |
| 17 | city | high | [17] contains city names (DUBOIS, OLANTA, CLEARFIELD, CURWENSVILLE) |
| 18 | state | high | [18] all values are 'PA' (Pennsylvania abbreviation) |
| 19 | zip | high | [19] contains valid 5-digit ZIP codes (15801, 16863, 16830, 16833) |
| 150 | skip | high | [150] contains 10-digit US phone number (8142363267) |
| 151 | city | high | [151] repeated city name 'CLEARFIELD' |
Notes: 153 total columns; 12 contain PII. Key PII fields identified: full names (lastName, firstName, middleName, suffix), gender, DOB (birth date), residential address components (address1, address2, city, state, zip), and phone number. Columns 0, 1, 8-13, 20-149 (except mapped) appear to be internal IDs, precinct codes, status flags, or empty — mapped to skip per rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CLEARFIELD__CLEARFIELD_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data appears to be structured geographic and administrative codes for Clearfield County, Pennsylvania, including boroughs, townships, legislative districts, and county identifiers. There are no visible PII fields (e.g., names, addresses, emails, phone numbers, DOB, SSN) in the first 50 rows. All entries consist of location names, district numbers, and codes, which do not qualify as personally identifiable information under the provided field definitions. No columns map to any PII categories.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CLEARFIELD__CLEARFIELD_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: yes·Quote: "
Notes: This is a metadata header file describing column meanings for voter data, not actual voter records. The rows shown are column descriptions rather than data rows. No PII fields are present in this header metadata file — it only defines geographic/political divisions (Precinct, Ward, School district, Municipality, Magistrate, Legislative, Senate, Congressional, County). These are administrative codes, not personal identifiers. Therefore, no columns map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CLINTON__CLINTON_Election_Map_20181001.txt0 rows
File structure
Notes: free-form text with no consistent delimiter or row structure. Contains only election-related text and no identifiable personal information columns
USVoterData_BF__data__Pennsylvania__2018__Statewide__CLINTON__CLINTON_FVE_20181001.txt12 columns20,736 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header implies last name, values are common surnames |
| 3 | firstName | high | [3] header implies first name, values are common given names |
| 4 | middleName | medium | [4] values include single letters and occasional full middle names |
| 6 | gender | high | [6] values are 'M'/'F' which map to gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, header suggests birth date |
| 14 | address1 | high | [14] values are street addresses |
| 16 | city | high | [16] values are city names |
| 17 | city | high | [17] duplicate city values |
| 18 | state | high | [18] values are US state abbreviations (PA) |
| 19 | zip | high | [19] values are 5-digit ZIP codes |
| 150 | skip | medium | [150] contains a 10-digit number matching US phone format |
| 151 | fullName | high | [151] values repeat 'CLINTON' suggesting a formal name field |
Notes: 153 total columns; 22 contain PII including names, addresses, DOB, gender, phone. Many columns appear to be internal state/voter IDs, precinct codes, and party affiliation flags which are skipped per rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CLINTON__CLINTON_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file appears to be a structured list of geographic and administrative divisions rather than a dataset containing personal identifiable information (PII). It includes entries for townships, boroughs, regions, districts, and county codes, but no actual voter records with names, addresses, or other PII fields. All entries are geographic identifiers and codes, which do not qualify as PII under the defined field types. Therefore, no columns map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CLINTON__CLINTON_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided sample contains only header-like structure and codes (CLINTON, Prec, WD, etc.) without any actual PII fields. These appear to be metadata columns describing geographic and administrative divisions (Precinct, Ward, School District, Municipality, etc.). No personal identifiable information such as names, addresses, dates of birth, phone numbers, etc., is present in this snippet. Therefore, no columns map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__COLUMBIA__COLUMBIA_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data appears to be a list of election events with no personal identifiable information (PII). The columns contain location (COLUMBIA), an index number, election description, and election date. There are no columns mapping to PII fields such as email, phone, DOB, names, addresses, etc. Therefore, no PII columns are identified.
USVoterData_BF__data__Pennsylvania__2018__Statewide__COLUMBIA__COLUMBIA_FVE_20181001.txt10 columns38,075 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header indicates last name, values are common surnames |
| 3 | firstName | high | [3] header indicates first name, values are common given names |
| 4 | middleName | medium | [4] values include single letters and full names, matches middle name pattern |
| 6 | gender | high | [6] values are 'M'/'F' which map to gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date format, header suggests birth date |
| 14 | address1 | high | [14] contains street addresses like 'CATHERINE ST', 'BLACK BEAR DR' |
| 16 | city | high | [16] contains city names like 'BLOOMSBURG', 'CATAWISSA' |
| 18 | state | high | [18] contains state abbreviation 'PA' |
| 19 | zip | high | [19] contains 5-digit ZIP codes |
| 151 | fullName | medium | [151] contains city name 'COLUMBIA' which could be part of full name in some contexts, but low confidence due to limited samples |
Notes: 153 total columns, 10 contain PII. Columns 0, 1, 5, 8, 9, 10, 11-153 contain internal IDs, timestamps, flags, or empty values and are skipped per rules. Column 151 maps to fullName with medium confidence due to limited samples but is included as it may represent part of a name in some records.
USVoterData_BF__data__Pennsylvania__2018__Statewide__COLUMBIA__COLUMBIA_Zone_Codes_20181001.txt0 rows
File structure
Notes: free-form text with no consistent column structure. Data consists of geographic references and district codes for Columbia County, PA, formatted as free text lines without a shared delimiter or repeating structure. No PII fields are present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__COLUMBIA__COLUMBIA_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file is a structured dataset with headers indicating geographic/political divisions (Precinct, Ward, School district, etc.). No columns map to PII fields (email, phone, dob, name, address, etc.). All visible columns are administrative codes or region identifiers, not personal data. No personal identifiers are present in the provided rows.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CRAWFORD__CRAWFORD_Election_Map_20181001.txt0 rows
File structure
Notes: This file is free-form text with no consistent columnar structure. Each line contains voter election participation records in prose format, not a delimited CSV. Since it lacks consistent delimiters and row structure, it qualifies as unstructured text requiring streaming extraction. No column mapping is possible.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CRAWFORD__CRAWFORD_FVE_20181001.txt13 columns52,223 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] values are common last names: NELSON, PRICE, ALDRETE, MILLER, LEWIS |
| 3 | firstName | high | [3] values are common first names: KATHERINE, JEFFREY, SUSAN, HENRY, BRIDGET, KAREN |
| 4 | middleName | high | [4] contains middle names and initials: 'A O', 'EVERETT', 'L', 'F', 'BJURSTROM', 'C' |
| 5 | suffix | high | [5] contains honorific suffix values: 'SR' |
| 6 | gender | high | [6] binary gender codes: 'F' and 'M' |
| 7 | skip | high | [7] dates in MM/DD/YYYY format likely birth dates: 09/04/1937, 09/26/1964, etc. |
| 8 | skip | high | [8] additional birth dates in MM/DD/YYYY format: 01/01/1972, 12/11/1997, etc. |
| 14 | address1 | high | [14] contains street addresses: 'BIRCH DR', 'LOCUST ST', 'EASTVIEW AVE', etc. |
| 17 | city | high | [17] contains city names: 'MEADVILLE', 'CAMBRIDGE SPRINGS' |
| 18 | state | high | [18] two-letter state abbreviation: 'PA' |
| 19 | zip | high | [19] contains 5-digit ZIP codes: '16335', '16403' |
| 150 | skip | low | [150] one value appears to be a 10-digit phone number: '8145734148' |
| 151 | fullName | high | [151] contains full name: 'CRAWFORD' (likely part of a full name construction from surrounding name fields) |
Notes: 153 columns total, 12 contain PII: lastName, firstName, middleName, suffix, gender, dob (two date columns), address1, city, state, zip, phone, and fullName. Remaining columns are internal IDs, geographic/district codes, status flags, and empty fields — all mapped to skip per exclusion rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CRAWFORD__CRAWFORD_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is not structured PII data. The file contains geographic and administrative codes for Crawford County, Pennsylvania, including township names, school districts, legislative districts, and congressional districts. There are no personal identifiers (names, addresses, DOBs, etc.) present. All entries are location-based codes and labels, which do not constitute PII under the defined field types.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CRAWFORD__CRAWFORD_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a structured file but contains only non-PII metadata columns (location codes, district identifiers). No personal identifying information fields (names, addresses, DOB, etc.) appear in the first 50 rows. All columns represent geographic/political divisions or placeholder values.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CUMBERLAND__CUMBERLAND_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is not structured voter registration data but a list of election events with precinct names and dates. No PII columns are present — only election names and dates. All columns are skipped per exclusion rules (no names, addresses, DOB, etc.).
USVoterData_BF__data__Pennsylvania__2018__Statewide__CUMBERLAND__CUMBERLAND_FVE_20181001.txt12 columns170,865 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] headerless column values are surnames (SIMMONS, REXROTH, AUX, NOPSON, LEEDS, BOWES) |
| 3 | firstName | high | [3] headerless column values are first names (RICHARD, MARTHA, JAMES, LINDA, FREDERICK, THOMAS) |
| 4 | middleName | medium | [4] single initial values (E, H, R, R, A, M) match typical middle name initials |
| 5 | suffix | high | [5] contains suffix values like 'JR' matching generational suffix pattern |
| 6 | gender | high | [6] contains gender codes 'M' and 'F' (male/female) |
| 7 | skip | high | [7] values match MM/DD/YYYY date format (05/14/1958, 12/08/1938, 12/26/1948, 11/02/1941, 04/21/1945, 11/10/1934) |
| 8 | skip | high | [8] additional date-of-birth values in MM/DD/YYYY format (01/01/1976 repeated across samples) |
| 14 | address1 | high | [14] contains street address values (CHARLES ST, SHIPPENSBURG RD, MOORELAND AVE, GEORGE AVE, LEEDS RD, BULLOCK CIR) |
| 16 | city | high | [16] contains city names (MECHANICSBURG, SHIPPENSBURG, CARLISLE, NEWVILLE) |
| 17 | state | high | [17] all values are 'PA' indicating Pennsylvania state |
| 18 | zip | high | [18] contains 5-digit ZIP codes (17055, 17257, 17013, 17241, 17015) |
| 151 | county | high | [151] contains repeated 'CUMBERLAND' values indicating county name |
Notes: 153 total columns identified; 12 contain PII (names, DOB, address, gender, state/ZIP). Columns 0,1,13,15,20-22,24-150 contain internal IDs, empty fields, or election-specific codes (precinct, district, party codes) — mapped to skip. Election party codes (R/D/AP) are not PII and excluded.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CUMBERLAND__CUMBERLAND_Zone_Codes_20181001.txt0 rows
File structure
Notes: The file contains purely geographic / administrative codes (precinct names, school districts, municipality codes, legislative districts) with no personal or identifiable information. All columns map to non-PII administrative categories per voter data schema guidelines. No email, phone, name, address, DOB, SSN, or other PII fields present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__CUMBERLAND__CUMBERLAND_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a metadata header file describing geographic districts and voting structures, not actual voter records. Columns contain geographic codes, district names, and structural identifiers — no personal PII fields (names, addresses, DOB, etc.) are present. All columns map to skip per exclusion rules (company/business references, internal IDs, and non-PII descriptors).
USVoterData_BF__data__Pennsylvania__2018__Statewide__DAUPHIN__DAUPHIN_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains election-related data with no personal identifiable information (PII). All columns appear to be structured election event records (election type, date) without any voter names, addresses, or other PII fields. The data represents election cycles rather than individual voter details.
USVoterData_BF__data__Pennsylvania__2018__Statewide__DAUPHIN__DAUPHIN_FVE_20181001.txt16 columns184,152 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header implies last name, values are common surnames |
| 3 | firstName | high | [3] header implies first name, values are common given names |
| 4 | suffix | high | [4] single letters likely generational suffixes (Jr/Sr), not honorifics |
| 6 | gender | high | [6] values are 'M'/'F' — unambiguous gender indicators |
| 7 | skip | high | [7] MM/DD/YYYY dates likely birthdate (earliest dates 1919-1946) |
| 8 | skip | high | [8] MM/DD/YYYY dates likely alternate birthdate or registration date; maps to dob per rules |
| 14 | address1 | high | [14] free-text street addresses like 'THREE RIVERS DR' |
| 15 | address2 | medium | [15] numeric values likely apartment/unit numbers |
| 16 | address2 | medium | [16] values like 'APT 404' clearly secondary address line |
| 17 | city | high | [17] city names like 'HARRISBURG', 'PALMYRA', 'MIDDLETOWN' |
| 18 | state | high | [18] consistent 'PA' values — state abbreviation |
| 19 | zip | high | [19] 5-digit zip codes like '17112', '17078' |
| 20 | address1 | medium | [20] contains additional street address fragments like '130 E LAUER LN' |
| 22 | city | medium | [22] contains city name 'CAMP HILL' |
| 23 | state | medium | [23] contains state abbreviation 'PA' |
| 24 | zip | medium | [24] contains 5-digit zip code '17011' |
Notes: 153 columns total, 22 contain PII. Excluded internal IDs (0, 12, 26, 27, 30, 31, 32-38, 42, etc.), timestamps (7,8,10,25,28), political/registration flags (11,89-149), and other non-PII fields per exclusion rules. Multiple address lines captured where present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__DAUPHIN__DAUPHIN_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data is purely administrative codes for Pennsylvania voting districts and precincts. It contains no personal identifiable information (PII). Columns represent geographic regions (counties, boroughs, townships) and electoral subdivisions, but no names, addresses, dates of birth, etc. No mapping to PII fields is possible.
USVoterData_BF__data__Pennsylvania__2018__Statewide__DAUPHIN__DAUPHIN_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: none
Notes: This appears to be a header/metadata row structure, not actual voter data records. The first row contains the literal string 'DAUPHIN' repeated, followed by column labels like 'Precinct', 'City ward', etc. These are administrative codes and geographic divisions, not PII. No actual voter records (names, addresses, DOB, etc.) are present in the first 15 rows provided, which are purely structural. Additional rows containing voter data would need to be examined to map PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__DELAWARE__DELAWARE_Election_Map_20181001.txt0 rows
File structure
Notes: free-form text with no consistent delimiter or row structure; contains election metadata but no PII fields
USVoterData_BF__data__Pennsylvania__2018__Statewide__DELAWARE__DELAWARE_FVE_20181001.txt13 columns399,153 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header implies last name, values are common surnames |
| 3 | firstName | high | [3] header implies first name, values are common given names |
| 4 | suffix | high | [4] values are single letters (R, W, G, M) which could be honorifics or suffixes; given the voter data context and lack of explicit title column, these are likely suffixes |
| 6 | gender | high | [6] values are 'F', 'M', 'U' which correspond to gender codes |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, typical for birth dates in voter data |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern, typical for birth dates in voter data |
| 11 | suffix | high | [11] values are party affiliations (R, D, CT), but given the pattern of other columns and voter data context, this is likely a mislabeled suffix column. However, party affiliation is not a PII field, so this should be skipped. Re-evaluating: In voter data, party affiliation is typically not considered PII. Therefore, this column should be skipped. |
| 14 | address1 | high | [14] values are street addresses |
| 15 | address2 | medium | [15] values like '3305', '11A', 'C9' suggest apartment/unit numbers |
| 17 | city | high | [17] values are city names |
| 18 | state | high | [18] values are all 'PA' indicating state |
| 19 | zip | high | [19] values are 5-digit zip codes |
| 151 | country | high | [151] values are all 'DELAWARE' indicating country/state |
Notes: 153 columns total, 11 contain PII. This is a US voter data compilation with detailed personal information including names, addresses, birth dates, gender, and state. Columns 4 and 11 were identified as potential suffix and party affiliation respectively, but party affiliation is not a PII field and should be skipped. All other columns appear to be internal IDs, geographic codes, or timestamps and are skipped per exclusion rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__DELAWARE__DELAWARE_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data is purely geographic identifiers (state, precinct code, precinct name) with no personal identifying information (PII). All columns represent administrative divisions and voting precincts, not individual voter records. No email, phone, name, address, DOB, SSN, etc. are present in the sample.
USVoterData_BF__data__Pennsylvania__2018__Statewide__DELAWARE__DELAWARE_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: yes·Quote: "
Notes: This is a metadata header file describing voter registration fields for Delaware, not actual voter records. It contains only geographic/district labels (Precinct, Ward, School district, Municipality, etc.) and administrative codes. No personal PII fields (names, addresses, DOB, etc.) appear in this sample. All columns represent internal codes and descriptions, so they map to 'skip' per rules. The file is structured but contains zero PII.
USVoterData_BF__data__Pennsylvania__2018__Statewide__ELK__ELK_Election_Map_20181001.txt0 rows
File structure
Notes: free-form text with no consistent delimiter structure; lines vary in format and length. This is not a structured CSV/TSV file and contains election event descriptions rather than voter PII records. No column mapping is possible.
USVoterData_BF__data__Pennsylvania__2018__Statewide__ELK__ELK_FVE_20181001.txt13 columns19,224 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header implies last name, values are common surnames |
| 3 | firstName | high | [3] header implies first name, values are common given names |
| 4 | middleName | medium | [4] occasional middle initials/names present, header position suggests middle name |
| 5 | suffix | medium | [5] values are generational suffixes (Sr, Jr, etc.) |
| 6 | gender | high | [6] values are 'M' indicating male gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern and represent birth dates |
| 8 | skip | high | [8] additional birth date column with MM/DD/YYYY pattern |
| 14 | address1 | high | [14] values are street addresses (e.g., 'JEFFERSON ST') |
| 16 | address2 | medium | [16] values include 'PO BOX 102' indicating secondary address line |
| 17 | city | high | [17] values are city names (e.g., 'BYRNEDALE') |
| 18 | state | high | [18] values are state abbreviations (PA) |
| 19 | zip | high | [19] values are 5-digit ZIP codes |
| 150 | skip | high | [150] values are 10-digit phone numbers |
Notes: 153 total columns; 12 contain PII (names, addresses, DOB, gender, phone). Remaining columns are internal IDs, voting districts, status codes, and other non-PII administrative data. All PII fields mapped with high confidence based on header names and value patterns.
USVoterData_BF__data__Pennsylvania__2018__Statewide__ELK__ELK_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: File is unstructured: contains only geographic/taxonomic codes and names of locations/districts (counties, townships, school districts, legislative districts). No personal PII fields detected. Format is tab-delimited but rows represent location codes rather than individual voter records.
USVoterData_BF__data__Pennsylvania__2018__Statewide__ELK__ELK_Zone_Types_20181001.txt0 rows
File structure
Notes: The file appears to be a structured table of voter registration metadata with codes and descriptions, but no actual PII fields (names, addresses, DOB, etc.) are present in the first 50 rows. All columns contain geographic/political codes and labels (Precinct, Ward, School district, etc.). No emails, phone numbers, names, or other PII are visible. Therefore, no PII columns can be mapped.
USVoterData_BF__data__Pennsylvania__2018__Statewide__ERIE__ERIE_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains only election metadata (county, election ID, election name, date) with no personal voter data. All columns are skip fields per exclusion rules (no PII present).
USVoterData_BF__data__Pennsylvania__2018__Statewide__ERIE__ERIE_FVE_20181001.txt12 columns190,184 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header suggests surname, values are common last names |
| 3 | firstName | high | [3] header suggests given name, values are common first names |
| 4 | middleName | medium | [4] values are single letters, likely middle initials |
| 5 | suffix | high | [5] header empty but values are generational suffixes (JR) |
| 6 | gender | high | [6] values are 'M'/'F' which map to gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, header implies birth date |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern, likely alternate birth date |
| 14 | address1 | high | [14] values are street addresses |
| 17 | city | high | [17] values are city names |
| 18 | state | high | [18] values are state abbreviations (PA) |
| 19 | zip | high | [19] values are 5-digit zip codes |
| 150 | skip | high | [150] values are 10-digit numbers matching US phone format |
Notes: 153 columns total, 12 contain PII. Columns 0,1,13,15-152 appear to be internal IDs, empty fields, or administrative codes and are skipped. Dates at indices 10,24,28 are registration/voter status dates — not DOB — skipped per rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__ERIE__ERIE_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The file contains purely geographic and administrative data (ward districts, school districts, municipalities) with no identifiable personal information. All columns represent location codes and names of districts/townships, which do not qualify as PII under the defined categories. No email addresses, phone numbers, names, or other personal data appear in the sample.
USVoterData_BF__data__Pennsylvania__2018__Statewide__ERIE__ERIE_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header-only file describing voter registration fields for the state of ERIE. No actual voter data (names, addresses, etc.) is present in the first 50 rows — only metadata about district types and codes. Since there are no PII values or structured records, no columns can be mapped to PII fields. The file appears to be a schema or field definition listing rather than actual voter records.
USVoterData_BF__data__Pennsylvania__2018__Statewide__FAYETTE__FAYETTE_Election_Map_20181001.txt0 rows
File structure
Notes: This file contains no structured columns. It is a list of election events with county names, event IDs, event descriptions, and dates. There are no personal identifiable information (PII) fields such as names, addresses, or contact details in the provided rows. The data appears to be election metadata rather than voter records.
USVoterData_BF__data__Pennsylvania__2018__Statewide__FAYETTE__FAYETTE_FVE_20181001.txt10 columns77,646 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header 'WOLFE', values are surnames |
| 3 | firstName | high | [3] header 'ELEANORE', values are given names |
| 4 | middleName | medium | [4] header 'MARIE', values are initials or middle names |
| 6 | gender | high | [6] values 'F'/'M' map directly to gender |
| 7 | skip | high | [7] values match MM/DD/YYYY birth date pattern |
| 14 | address1 | high | [14] header 'HOPEWELL RD', values are street addresses |
| 17 | city | high | [17] header 'WHITE', values are city names |
| 18 | state | high | [18] values are consistent US state abbreviation 'PA' |
| 19 | zip | high | [19] 5-digit values represent ZIP codes |
| 150 | skip | medium | [150] values are 10-digit US phone numbers |
Notes: 153 columns total, 10 contain PII: lastName, firstName, middleName, gender, dob, address1, city, state, zip, phone. All other columns are internal IDs, geographic codes, voter status flags, or empty fields and are skipped per rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__FAYETTE__FAYETTE_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data consists entirely of geographic and administrative codes (county, school districts, municipalities, judicial districts, legislative districts) with no personal identifying information (PII). There are no columns containing names, addresses, dates of birth, phone numbers, emails, or other PII fields. All entries are location identifiers and district codes, which fall under skip criteria (internal IDs, administrative codes).
USVoterData_BF__data__Pennsylvania__2018__Statewide__FAYETTE__FAYETTE_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header row describing geographic/district codes (Precinct, Ward, School District, etc.) and not actual voter data rows. No PII fields present in this row. The file appears to be a metadata header for a voter registration dataset; actual PII-containing rows would follow this pattern but aren't visible in the first 50 rows provided.
USVoterData_BF__data__Pennsylvania__2018__Statewide__FOREST__FOREST_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is not structured PII data. The file contains election-related metadata: election titles, sequence numbers, and election dates. No personal identifiable information (PII) such as names, addresses, or contact details is present. The data appears to be a list of election events rather than voter records.
USVoterData_BF__data__Pennsylvania__2018__Statewide__FOREST__FOREST_FVE_20181001.txt14 columns3,319 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header implies last name, values are surnames |
| 3 | firstName | high | [3] header implies first name, values are common given names |
| 4 | middleName | medium | [4] headerless but values match typical middle name patterns (single letters, short names) |
| 5 | suffix | high | [5] values are generational suffixes (SR, III, JR) |
| 6 | gender | high | [6] values are M/F, headerless but clearly gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, likely birthdate |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern, likely registration or election date — but given context and adjacent birth date column, this is also a date of birth component; map as dob |
| 10 | skip | high | [10] values match MM/DD/YYYY date pattern, likely registration date — but given context and adjacent birth date columns, this may also be a date of birth component; map as dob |
| 14 | address1 | high | [14] values are street addresses |
| 16 | address2 | medium | [16] values are PO Boxes, secondary address lines |
| 17 | city | high | [17] values are city names |
| 18 | state | high | [18] values are state abbreviations (PA) |
| 19 | zip | high | [19] values are 5-digit zip codes |
| 151 | fullName | high | [151] values repeat 'FOREST' — likely a placeholder or error, but if real would be a full name component; map as fullName based on position and context |
Notes: 153 columns total, 17 contain PII (names, addresses, DOB, gender, suffix). Many columns are internal IDs, precinct codes, party affiliation, and other non-PII metadata. Columns 7, 8, and 10 all map to dob due to date patterns and voter data context — while 8 and 10 may represent registration or election dates, they follow date formats consistent with birth dates and are adjacent to a clear dob column, justifying the mapping.
USVoterData_BF__data__Pennsylvania__2018__Statewide__FOREST__FOREST_Zone_Codes_20181001.txt0 rows
File structure
Notes: This file contains only geographic and administrative codes (townships, districts, counties) with no personal identifiable information. All columns appear to be location codes, district identifiers, and township names which do not contain any PII fields. No email, phone, name, address, or other PII fields are present in the sample data.
USVoterData_BF__data__Pennsylvania__2018__Statewide__FOREST__FOREST_Zone_Types_20181001.txt0 rows
File structure
Notes: This file appears to be a header-less metadata file describing voting districts and administrative divisions, not actual voter records. The columns contain geographic/district codes (Precinct, Ward, School district, Municipality, Magistrate district, Legislative, Senate, Congressional, County) and "Not used" placeholders. No PII fields (names, addresses, DOB, etc.) are present in the first 50 rows. These appear to be structural codes for organizing voter rolls rather than personal voter data.
USVoterData_BF__data__Pennsylvania__2018__Statewide__FRANKLIN__FRANKLIN_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains only election-related data (voter precinct names, election types, and dates). No PII fields (names, addresses, DOB, etc.) are present. All entries are structured records of election events with no personal identifying information. Therefore, no columns map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__FRANKLIN__FRANKLIN_FVE_20181001.txt12 columns91,957 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header implies surname, values are common last names |
| 3 | firstName | high | [3] header implies given name, values are common first names |
| 4 | middleName | medium | [4] values include both single initials and full middle names |
| 6 | gender | high | [6] values are 'M'/'F' which map directly to gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern and represent birth dates |
| 8 | skip | high | [8] values match YYYY-MM-DD date pattern and represent birth dates |
| 14 | address1 | high | [14] contains street address components like 'ORRSTOWN RD' and 'MENNO VILLAGE' |
| 17 | city | high | [17] values are city names like 'ORRSTOWN', 'CHAMBERSBURG', 'SHIPPENSBURG' |
| 18 | state | high | [18] values are all US state abbreviations ('PA') |
| 19 | zip | high | [19] values are 5-digit US zip codes |
| 150 | skip | high | [150] values are 10-digit numbers matching US phone format |
| 151 | fullName | high | [151] values are consistent surname 'FRANKLIN' repeated across rows, likely part of full name construction |
Notes: 153 total columns identified; 11 contain PII (names, addresses, dates of birth, phone numbers, gender, state, zip). Remaining columns are internal IDs, geographic codes, status flags, and empty fields — all mapped to skip per exclusion rules. Data structure matches US voter registration format with detailed personal and precinct information.
USVoterData_BF__data__Pennsylvania__2018__Statewide__FRANKLIN__FRANKLIN_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The file is purely structural/geographic data (precinct, district, and legislative district codes) with no personal or identifiable information. All columns represent location identifiers, district numbers, and township/borough names — no PII fields present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__FRANKLIN__FRANKLIN_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header-less metadata file describing district/precinct codes rather than containing actual voter records. No PII fields present. All columns contain district codes, labels, and structural descriptors (e.g., 'School district', 'Municipal', 'Judicial').
USVoterData_BF__data__Pennsylvania__2018__Statewide__FULTON__FULTON_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file appears to be a list of election records with no personal identifiable information (PII). The columns contain election names and dates, which do not map to any PII fields. Therefore, no columns are mapped to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__FULTON__FULTON_FVE_20181001.txt12 columns9,025 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header implies last name, values are common surnames |
| 3 | firstName | high | [3] header implies first name, values are common given names |
| 4 | middleName | high | [4] header implies middle name, values are single letters representing initials |
| 6 | gender | high | [6] values are 'M'/'F' which map directly to gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern and represent birth dates |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern and represent birth dates |
| 14 | address1 | high | [14] values are street addresses |
| 15 | address2 | high | [15] values are apartment/suite numbers |
| 16 | city | high | [16] values are city names |
| 18 | state | high | [18] values are US state abbreviations |
| 19 | zip | high | [19] values are 5-digit ZIP codes |
| 150 | skip | high | [150] values are 10-digit US phone numbers |
Notes: 153 columns total, 12 contain PII. Columns 0,1,5,9-13,20-23,25-32,34-152 contain internal IDs, timestamps, or geographic codes that are not PII. The data represents US voter registrations with detailed personal information.
USVoterData_BF__data__Pennsylvania__2018__Statewide__FULTON__FULTON_Zone_Codes_20181001.txt0 rows
File structure
Notes: This is a structured dataset of US voter registration data, but the provided rows are not actual voter records—they are metadata describing townships, school districts, magisterial districts, legislative districts, congressional districts, and county references. The rows contain location identifiers (e.g., "AYR TOWNSHIP", "CENTRAL FULTON SCHOOL DISTRICT"), codes (e.g., "0100", "SC29130", "MN001"), and district names, but no personal voter information such as names, addresses, DOBs, or contact details. Therefore, there are no PII fields in these rows to map. The file likely contains additional structured voter records elsewhere; this snippet only shows geographic/district metadata.
USVoterData_BF__data__Pennsylvania__2018__Statewide__FULTON__FULTON_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a metadata header row for precinct/district codes, not actual voter records. No PII present. All columns describe geographic/district codes and are non-PII per exclusion rules (internal codes, district identifiers). No email, phone, name, address, or other PII fields appear in this row.
USVoterData_BF__data__Pennsylvania__2018__Statewide__GREENE__GREENE_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a simple two-column tab-delimited file with election metadata only. Column 0 appears to be a county name ("GREENE") and column 1 contains sequential record numbers. Column 2 contains election names and column 3 contains election dates. There are no PII fields present in this data sample.
USVoterData_BF__data__Pennsylvania__2018__Statewide__GREENE__GREENE_FVE_20181001.txt14 columns21,762 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header implies last name, values are common surnames |
| 3 | firstName | high | [3] header implies first name, values are common given names |
| 4 | middleName | medium | [4] headerless but values match middle names/initials |
| 6 | gender | high | [6] values are 'M'/'F' and header implies gender |
| 7 | skip | high | [7] values match MM/DD/YYYY birth date format |
| 14 | address1 | high | [14] values are street addresses |
| 17 | city | high | [17] values are city names |
| 18 | state | high | [18] values are two-letter US state abbreviations |
| 19 | zip | high | [19] values are 5-digit US zip codes |
| 20 | address2 | medium | [20] values include PO boxes and addresses |
| 22 | city | medium | [22] duplicate city values |
| 23 | state | medium | [23] duplicate state values |
| 24 | zip | medium | [24] duplicate zip values |
| 150 | skip | high | [150] values are 10-digit US phone numbers |
Notes: 153 columns total, 13 contain PII. Columns 0,1,5,8-13,15,21,25-152 contain internal IDs, timestamps, flags, or district codes and are skipped. Column 8 appears to be an alternate birth date but is excluded due to headerless nature and overlapping meaning with column 7.
USVoterData_BF__data__Pennsylvania__2018__Statewide__GREENE__GREENE_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data is purely geographic and administrative codes with no personal identifiable information (PII). Columns contain county names, district codes, and municipal/township identifiers — none map to email, phone, name, address, DOB, SSN, gender, suffix, or other PII fields. No structured PII exists in this sample.
USVoterData_BF__data__Pennsylvania__2018__Statewide__GREENE__GREENE_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a metadata file describing district codes and precinct structures for Greene County. It contains no personal voter records or PII fields - only geographic identifiers and administrative codes. All values are codes and descriptions (e.g., 'Precinct', 'Ward', 'School district'). No columns map to PII fields as per the exclusion rules for internal IDs and non-PII codes.
USVoterData_BF__data__Pennsylvania__2018__Statewide__HUNTINGDON__HUNTINGDON_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is not structured PII data; it's a list of election events with place names, identifiers, and dates. No columns map to PII fields. The data appears to be election metadata rather than voter registration records.
USVoterData_BF__data__Pennsylvania__2018__Statewide__HUNTINGDON__HUNTINGDON_FVE_20181001.txt12 columns29,868 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header 'lastName' inferred from values like DUNLAVY, SEAGRIST, BUKOWSKI |
| 3 | firstName | high | [3] header 'firstName' inferred from values like MARION, LARRY, TRUDIE |
| 4 | middleName | medium | [4] values include single letters and names like MAE, JOSEPH |
| 6 | gender | high | [6] values are 'F'/'M' which map to gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern and represent birth dates |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern and represent birth dates |
| 14 | address1 | high | [14] contains street address components like WESTMINSTER DR, SHADE VALLEY RD |
| 17 | city | high | [17] contains city names like HUNTINGDON, SHADE GAP, JAMES CREEK |
| 18 | state | high | [18] all values are 'PA' indicating Pennsylvania |
| 19 | zip | high | [19] contains 5-digit ZIP codes like 16652, 17255, 16657 |
| 150 | skip | high | [150] contains 10-digit US phone numbers |
| 151 | city | high | [151] duplicate city values like HUNTINGDON |
Notes: 153 total columns; 12 contain PII (names, dob, gender, address, phone). Columns 0,1,5,13,15-152 are internal IDs, empty fields, or administrative codes and skipped per rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__HUNTINGDON__HUNTINGDON_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data consists of structured geographical and administrative codes without any personal identifiable information (PII). All fields appear to represent location codes, district identifiers, and region labels rather than individual voter data. No columns map to PII fields such as names, addresses, emails, or phone numbers.
USVoterData_BF__data__Pennsylvania__2018__Statewide__HUNTINGDON__HUNTINGDON_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a metadata header row from a voter registration dataset. The first row describes column purposes (e.g., Precinct, City ward, School district) but contains no actual voter PII. All columns here are structural identifiers (Precinct, WD, SC, etc.) and 'Not used' placeholders — none map to PII fields. True PII (names, addresses, DOB, etc.) will appear in subsequent rows.
USVoterData_BF__data__Pennsylvania__2018__Statewide__INDIANA__INDIANA_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains no PII columns. It appears to be a list of election types and dates for Indiana, with no personal identifying information present. All columns are election-related metadata (state, election ID, election description, election date) and do not contain any names, addresses, dates of birth, phone numbers, or other PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__INDIANA__INDIANA_FVE_20181001.txt12 columns48,888 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header implies last name, values are surnames (WHITACRE, BENFIELD, HENDRICKS, etc.) |
| 3 | firstName | high | [3] header implies first name, values are given names (J, ANGELA, DAVID, etc.) |
| 4 | middleName | high | [4] header implies middle name, values are single letters or initials (L, R, H, W, B) |
| 6 | gender | high | [6] values are 'F'/'M' matching gender codes, header implies gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date format, header implies birth date |
| 8 | skip | high | [8] values match MM/DD/YYYY date format, likely alternate birth date or registration date |
| 14 | address1 | high | [14] values are street addresses (ST ANDREWS CT, SNOW DRIFT LN, etc.) |
| 17 | state | high | [17] values are state names (INDIANA), header implies state |
| 18 | state | high | [18] values are state abbreviations (PA), header implies state |
| 19 | zip | high | [19] values are 5-digit ZIP codes (15701) |
| 150 | skip | high | [150] values are 10-digit US phone numbers (7244641803) |
| 151 | state | high | [151] values are state names (INDIANA), duplicate state info |
Notes: 153 total columns; 22 contain PII (names, DOB, gender, address, state, ZIP, phone). Remaining columns are internal IDs, precinct codes, voter status flags, and other non-PII administrative data. Data represents US voter registrations from multiple states.
USVoterData_BF__data__Pennsylvania__2018__Statewide__INDIANA__INDIANA_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a structured tabular file but contains no personal identifiable information (PII). All columns represent geographic/political divisions (state, district codes, and district names) with no names, addresses, contact details, or other PII fields. The data appears to be purely administrative geographic identifiers for Indiana voting districts and school districts.
USVoterData_BF__data__Pennsylvania__2018__Statewide__INDIANA__INDIANA_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header/metadata row describing columns for Indiana voter data. No actual PII fields are present in the first 50 rows. The rows contain state codes, column indices, and descriptions of what each column represents (e.g., precinct, ward, school district). These are structural identifiers, not personal data. Subsequent rows likely contain voter PII, but they are not visible in the provided sample.
USVoterData_BF__data__Pennsylvania__2018__Statewide__JEFFERSON__JEFFERSON_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a structured CSV file but contains no PII fields. The data appears to be election event metadata (election names and dates) with no personal information such as names, addresses, or voter details. All columns represent election identifiers and dates, which are not PII under the defined rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__JEFFERSON__JEFFERSON_FVE_20181001.txt10 columns29,626 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header implies last name, values are common surnames |
| 3 | firstName | high | [3] header implies first name, values are common given names |
| 4 | middleName | medium | [4] values are single letters and short names, consistent with middle names |
| 6 | gender | high | [6] values are 'M'/'F'/'U', maps directly to gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, labeled as birth date |
| 14 | address1 | high | [14] header indicates address, values are street addresses |
| 17 | city | high | [17] header and values clearly indicate city names |
| 18 | state | high | [18] values are state abbreviations (PA), maps to state |
| 19 | zip | high | [19] values are 5-digit zip codes |
| 150 | skip | high | [150] values are 10-digit phone numbers |
Notes: 153 columns total, 10 contain PII: names, gender, DOB, address components, phone. Remaining columns are internal IDs, voting districts, status flags, and empty fields — all skipped per rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__JEFFERSON__JEFFERSON_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains structured geographic and administrative region codes for Jefferson County. It lists townships, boroughs, districts, and regions with no personal identifying information (PII) present. All entries are place names, codes, and numerical identifiers with no names, addresses, dates of birth, contact details, or other PII fields. No columns map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__JEFFERSON__JEFFERSON_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header-only file describing column meanings for a voter registration dataset. No actual voter records are present in the first 50 rows — only placeholder labels like 'JEFFERSON' and descriptions such as 'Precinct', 'City ward', etc. Since there are no PII values and no structured data rows, no columns can be mapped to PII fields. All entries are metadata about geographic/district codes.
USVoterData_BF__data__Pennsylvania__2018__Statewide__JUNIATA__JUNIATA_Election_Map_20181001.txt0 rows
File structure
Notes: The provided data is free-form text listing election events with no consistent columnar structure. Each line describes an election with location, election type, and date, but there are no repeating fields like names, addresses, or other PII. This is not a delimited dataset and cannot be mapped to PII columns.
USVoterData_BF__data__Pennsylvania__2018__Statewide__JUNIATA__JUNIATA_FVE_20181001.txt14 columns13,732 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header 'WILLI', values are common surnames |
| 3 | firstName | high | [3] header 'AUSTIN', values are common first names |
| 4 | middleName | high | [4] header 'H', values like 'ANN' indicate middle names |
| 5 | suffix | high | [5] header 'JR', values 'JR', 'SR' are generational suffixes |
| 6 | gender | high | [6] header 'M', values 'M'/'F' represent gender |
| 7 | skip | high | [7] header '01/05/1947', values are dates in MM/DD/YYYY format |
| 8 | skip | high | [8] values are dates in MM/DD/YYYY format |
| 14 | address1 | high | [14] header 'NEALMAN RD', values are street addresses |
| 16 | city | high | [16] header 'MIFFLIN', values are city names |
| 17 | city | high | [17] values match known city names |
| 18 | state | high | [18] header 'PA', values are two-letter state abbreviations |
| 19 | zip | high | [19] header '17058', values are 5-digit ZIP codes |
| 150 | skip | medium | [150] values match 10-digit phone number pattern |
| 151 | state | high | [151] header 'JUNIATA', values are state names |
Notes: 153 total columns, 13 contain PII. Columns 0, 1, 9-13, 15, 20-149 (excluding mapped) are internal IDs, timestamps, flags, or empty. File is a voter registration dataset from US states.
USVoterData_BF__data__Pennsylvania__2018__Statewide__JUNIATA__JUNIATA_Zone_Codes_20181001.txt0 rows
File structure
Notes: The file contains purely geographic and administrative codes (county names, district numbers, region identifiers) with no personal identifiable information. All columns represent election districts, magisterial regions, legislative districts, and congressional codes — not voter records with PII. Therefore, no columns map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__JUNIATA__JUNIATA_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header descriptor file showing column mappings for a voter data CSV, not actual voter records. The lines describe field meanings (e.g., 'JUNIATA' is a county, 'Prec' is short for Precinct). No PII is present in this sample; it serves as a schema guide. All columns are metadata codes or placeholders (e.g., NU = Not Used). Therefore, no PII fields can be mapped from this structure definition alone.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LACKAWANNA__LACKAWANNA_Election_Map_20181001.txt0 rows
File structure
Notes: This is a structured file but contains only election event metadata (county name, event number, event description, date). No PII fields are present. All columns are election-related identifiers and timestamps, which are not considered PII under the provided rules. Therefore, no columns map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LACKAWANNA__LACKAWANNA_FVE_20181001.txt14 columns142,760 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header implies last name, values are common surnames |
| 3 | firstName | high | [3] header implies first name, values are common given names |
| 4 | middleName | high | [4] header implies middle name, values include initials and full middle names |
| 6 | gender | high | [6] values are 'M' and 'F', header suggests gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, header suggests birth date |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern, another birth date column |
| 14 | address1 | high | [14] values are street addresses, header suggests address |
| 17 | city | high | [17] values are city names, header suggests city |
| 18 | state | high | [18] values are state abbreviations, header suggests state |
| 19 | zip | high | [19] values are 5-digit zip codes, header suggests zip |
| 22 | city | medium | [22] values are city names, though header is empty |
| 23 | state | medium | [23] values are state abbreviations, though header is empty |
| 24 | zip | medium | [24] values are 5-digit zip codes, though header is empty |
| 151 | county | high | [151] values are county names, though no dedicated PII field exists; this is administrative location and not PII per se, but included for completeness in voter data context |
Notes: 153 columns total, 13 contain PII. This is a US voter registration dataset with detailed personal information including names, dates of birth, gender, and full residential addresses. Columns 7 and 8 both represent dates of birth (DOB), likely capturing different birth dates or registration dates. Column 151 contains county names, which are administrative locations and not direct PII but are included for context in voter data.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LACKAWANNA__LACKAWANNA_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is NOT a PII-containing structured file. The data consists entirely of geographic codes, municipality names, and school district identifiers. There are no personal identifiers, contact details, or sensitive attributes present in the visible rows. The format appears to be a tab-delimited list of voting precinct codes and associated administrative regions, typical of voter registration metadata files rather than individual voter records.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LACKAWANNA__LACKAWANNA_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header/metadata row structure describing voter district codes and abbreviations, not actual voter records. No PII fields present — all columns represent geographic/administrative codes (Precinct, Ward, School District, Municipality, etc.). The first column appears to be a location identifier ("LACKAWANNA"). This is typical of voter roll metadata files used to map geographic regions to internal codes; actual voter data would be in separate files with columns for names, addresses, DOB, etc.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LANCASTER__LANCASTER_Election_Map_20181001.txt0 rows
File structure
Notes: The provided text is a free-form list of election events with no consistent delimiter structure. Each line contains a location (LANCASTER), a sequential number, and election details, but there is no recurring column pattern or field separation. This format does not qualify as structured data with columns; it is purely descriptive election records. No PII fields are present in this sample.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LANCASTER__LANCASTER_FVE_20181001.txt13 columns325,172 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header is empty but values are surnames like GARBERICH, LEVERENTZ, HALL, EGAN, ZEIDERS, DARHOWER |
| 3 | firstName | high | [3] header is empty but values are first names like RICHARD, SHIRLEY, WILLIAM, JEAN, DELVIN, CHARLENE |
| 4 | suffix | high | [4] values are gender codes (R, C, L, H, E) but this column has values like 'R' and 'C' which are generational suffixes (Jr, Sr, etc.) in voter data |
| 5 | suffix | high | [5] values are 'SR' and empty, consistent with generational suffix (Sr.) in voter records |
| 6 | gender | high | [6] values are 'M' and 'F', clear gender indicators |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern and represent birth dates (07/24/1926, 05/22/1930, etc.) |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern and represent birth dates (01/01/1947, 01/01/1952, etc.) |
| 14 | address1 | high | [14] values are street addresses like FREEMASON DR, WATERCRESS LN, SYCAMORE DR, JAMES BUCHANAN DR, EDEN VIEW RD, W RIDGE RD |
| 17 | city | high | [17] all values are ELIZABETHTOWN, a clear city name |
| 18 | state | high | [18] all values are PA, a valid US state abbreviation |
| 19 | zip | high | [19] values are 5-digit zip codes like 17022 |
| 150 | skip | high | [150] values are 10-digit numbers (7173614094, 7173615799, etc.), typical US phone format |
| 151 | city | high | [151] all values are LANCASTER, a valid city name |
Notes: 153 columns total, 12 contain PII: lastName, firstName, suffix (2 columns), gender, dob (2 columns), address1, city, state, zip, phone. Remaining columns are internal IDs, precinct codes, registration dates, status flags, and other non-PII administrative data common in voter rolls.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LANCASTER__LANCASTER_Zone_Codes_20181001.txt0 rows
File structure
Notes: This file is a plain text list of geographic identifiers (precincts/townships) for Lancaster County, PA. It contains no personal identifiable information (PII) such as names, addresses, emails, or phone numbers. The structure is free-form with no consistent column arrangement. The entries are administrative divisions and do not contain any importable data for PII extraction.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LANCASTER__LANCASTER_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is not PII data; it's purely geographic/district metadata for Lancaster County. All columns describe political/administrative boundaries (precincts, wards, districts, etc.). No personal voter information appears in these rows. This appears to be a header/metadata row block from the voter file, not actual voter records.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LAWRENCE__LAWRENCE_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file appears to be a list of election events with no personal identifying information. The columns contain static text describing election types and dates, with no PII fields present. All rows follow the same pattern: location name, sequential number, event description, and date.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LAWRENCE__LAWRENCE_FVE_20181001.txt11 columns54,336 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] values are common last names: BOOK, BURCHETT, DEJOHN, BRUCKNER, GRECO |
| 3 | firstName | high | [3] values are common first names: RICHARD, RUSS, PHYLLIS, CHARLES, JEAN |
| 4 | middleName | medium | [4] values include initials and names: JAY, E, M, ADAM, L |
| 6 | gender | high | [6] values are M/F/U (male/female/undisclosed) |
| 7 | skip | high | [7] values match MM/DD/YYYY date format, typical for birth dates |
| 14 | address1 | high | [14] contains street addresses: HARBOR EDINBURG RD, SHENANGO RD, DEWEY AVE |
| 17 | city | high | [17] contains city names: EDINBURG, NEW CASTLE |
| 18 | state | high | [18] all values are PA (Pennsylvania) |
| 19 | zip | high | [19] contains 5-digit zip codes: 16116, 16105, 16101 |
| 150 | skip | low | [150] contains a 10-digit number: 7246987762 |
| 151 | fullName | high | [151] all values are LAWRENCE — likely a preserved full name field or family name |
Notes: 153 columns total, 11 contain PII. Columns 0, 1, 5, 8-13, 15-23, 25-62, 64-150 appear to be internal IDs, geographic codes, timestamps, or flags and are skipped per rules. Column 151 appears to be a preserved full name or family name field.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LAWRENCE__LAWRENCE_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data is purely geographic and administrative — it contains only place names, codes, and district identifiers with no personal identifying information (PII). No columns map to email, phone, dob, names, addresses, ssn, password, username, gender, suffix, or facebookId. All entries are structured as location references (townships, boroughs, precincts, school districts, magisterial districts, legislative districts, congressional districts, and county codes). Therefore, there are no PII columns to map.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LAWRENCE__LAWRENCE_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: yes·Quote: "
Notes: This file is a header-only or metadata file describing voter districts and codes, not actual voter records. The sample shows only geographic codes and labels with no personal data or PII fields. All columns are non-PII (geographic codes, precinct names, etc.).
USVoterData_BF__data__Pennsylvania__2018__Statewide__LEBANON__LEBANON_Election_Map_20181001.txt0 rows
File structure
Notes: This file is a list of election events with no personal voter data. It contains only election names, dates, and identifiers. No PII is present, and the structure is not tabular with consistent columns. The format is unstructured prose with no embedded emails or phone numbers.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LEBANON__LEBANON_FVE_20181001.txt12 columns85,026 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header suggests last name, values are common surnames: BEDARD, FACKLER, MOYLE, DELUCA, HONEYCHURCH, GREINER |
| 3 | firstName | high | [3] header suggests first name, values are common given names: ALBERT, MARY, JANE, VIVIAN, JOANNE, MARY |
| 4 | suffix | high | [4] header suggests suffix, values are generational suffixes: J, H, M, A, K |
| 6 | gender | high | [6] header suggests gender, values are M/F/U: M, F, F, F, U, F |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, header suggests birth date: 11/18/1928, 12/13/1924, 04/27/1924, 09/14/1929, 03/04/1938, 12/03/1925 |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern, header suggests birth date: 01/01/1956, 01/01/1958, 01/01/1976, 01/01/1980, 01/01/1968, 01/01/1968 |
| 14 | address1 | high | [14] header suggests address, values are street addresses: SWEETWATER DR, GRUBB RD, OXFORD DR, N LARKSPUR DR, CORNWALL MANOR , E MAIN ST |
| 17 | city | high | [17] header suggests city, values are city names: PALMYRA, LEBANON, CORNWALL, ANNVILLE |
| 18 | state | high | [18] header suggests state, values are state abbreviations: PA, PA, PA, PA, PA |
| 19 | zip | high | [19] header suggests zip, values are 5-digit zip codes: 17078, 17042, 17078, 17016, 17003 |
| 150 | skip | high | [150] values are 10-digit phone numbers: 7174504594, 7175667205 |
| 151 | city | high | [151] duplicate city values: LEBANON, LEBANON, LEBANON, LEBANON, LEBANON |
Notes: 153 columns total, 12 contain PII: lastName, firstName, suffix, gender, dob (two date columns), address1, city, state, zip, phone, duplicate city. All other columns are internal IDs, political affiliation codes, district/precinct identifiers, or empty fields — skipped per rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LEBANON__LEBANON_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
Notes: The provided sample data is purely administrative and geographic, listing political districts, school districts, and legislative districts for Lebanon County. No PII fields (email, phone, name, address, etc.) are present. All values are codes, names of municipalities, and district identifiers. Therefore, no columns map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LEBANON__LEBANON_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a metadata header file describing column purposes, not actual voter records. No PII present in this sample — only geographic codes, district labels, and placeholder 'Not used' entries. Actual voter data would contain PII in later rows, but this header-only snippet lacks any personal information.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LEHIGH__LEHIGH_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a list of election events with no PII fields. Columns contain election names and dates only. No identifiable personal information is present in the first 50 rows.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LEHIGH__LEHIGH_FVE_20181001.txt10 columns230,381 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header implies last name, values are common surnames |
| 3 | firstName | high | [3] header implies first name, values are common given names |
| 4 | middleName | high | [4] header implies middle name, values are single letters representing initials |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, labeled as birth date |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern, likely alternate birth date |
| 14 | address1 | high | [14] values are street addresses, labeled as street |
| 17 | city | high | [17] values are city names |
| 18 | state | high | [18] values are state abbreviations (PA) |
| 19 | zip | high | [19] values are 5-digit zip codes |
| 150 | skip | high | [150] values are 10-digit numbers, likely phone numbers |
Notes: 153 columns total, 10 contain PII: lastName, firstName, middleName, dob (two date columns), address1, city, state, zip, phone. Remaining columns are internal IDs, geographic codes, party affiliation, status flags, and other non-PII administrative data.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LEHIGH__LEHIGH_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is NOT a PII-containing dataset. The file consists entirely of geographic identifiers and codes (ward numbers, district codes, municipality names) with no personal identifying information such as names, addresses, dates of birth, or contact details. All values are administrative codes and location descriptors. No columns map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LEHIGH__LEHIGH_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file is a header-only or metadata file describing voter registration fields for Lehigh County. It contains no actual voter records or PII data. The lines are structured as [County] [Column ID] [Column Code] [Description], but no personal information is present in the first 50 rows provided. All entries describe geographic/political divisions (Precinct, Ward, School district, Municipality, etc.) and are not PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LUZERNE__LUZERNE_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The file contains election data with no PII fields. It lists election types and dates for Luzerne County, Pennsylvania, but does not include any personal identifiable information such as names, addresses, or contact details. All columns are skipped per exclusion rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LUZERNE__LUZERNE_FVE_20181001.txt11 columns206,080 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header 'MILGRIM' suggests lastName, values are surnames |
| 3 | firstName | high | [3] header 'NORMAN' suggests firstName, values are given names |
| 4 | middleName | medium | [4] contains mixed single letters and names; maps to middleName based on position and pattern |
| 6 | gender | high | [6] values 'M'/'F' map directly to gender |
| 7 | skip | high | [7] dates in MM/DD/YYYY format represent date of birth |
| 8 | skip | high | [8] additional date column also in MM/DD/YYYY format |
| 14 | address1 | high | [14] contains street addresses like 'EAST END BLVD' |
| 17 | city | high | [17] contains city names like 'WILKES BARRE' |
| 18 | state | high | [18] consistent 'PA' values indicate state |
| 19 | zip | high | [19] 5-digit postal codes |
| 150 | skip | high | [150] contains 10-digit phone numbers |
Notes: 153 total columns; 11 contain PII (names, addresses, dates, gender, phone). Remaining columns are internal IDs, geographic codes, election data, and empty fields — all skipped per rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LUZERNE__LUZERNE_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data appears to be structured but contains no identifiable PII fields. The columns appear to represent geographic and administrative divisions (e.g., townships, boroughs, wards) with codes and area names. There are no columns containing names, addresses, dates of birth, phone numbers, emails, SSNs, or other PII as defined in the instructions. All visible data is geographic/administrative categorization.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LUZERNE__LUZERNE_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header-only row containing geographic and administrative codes (Precinct, Ward, School district, Municipality, etc.). No PII fields present in this row. The data appears to be a structural metadata row defining codes for voter registration data rather than actual voter records. Subsequent rows would contain voter data, but this sample only includes the header row.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LYCOMING__LYCOMING_Election_Map_20181001.txt0 rows
File structure
Notes: The provided data appears to be a list of election events with no consistent column structure or personal information. Each line contains a county name, an entry number, an election description, and a date. There are no identifiable PII fields such as names, addresses, emails, or phone numbers in this sample. This is likely a reference or index file rather than a voter registration record dump.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LYCOMING__LYCOMING_FVE_20181001.txt12 columns67,843 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header 'lastName' inferred from sample values (ORNDORF, YOUNG, HARRIS, etc.), mapped as lastName |
| 3 | firstName | high | [3] header 'firstName' inferred from sample values (JANICE, ROBERT, MELISSA, etc.), mapped as firstName |
| 4 | middleName | high | [4] header 'middleName' inferred from sample values (A, KAY, T, LYNN, A), mapped as middleName |
| 6 | gender | high | [6] header 'gender' inferred from sample values (F, M, F, U, F, M), mapped as gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, header suggests birth date, mapped as dob |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern, header suggests birth date, mapped as dob |
| 14 | address1 | high | [14] header 'address1' inferred from sample values (RAMSEY DR, NORTHWAY ROAD EXT, BACK RD, etc.), mapped as address1 |
| 17 | city | high | [17] header 'city' inferred from sample values (JERSEY SHORE, WILLIAMSPORT, ALLENWOOD, etc.), mapped as city |
| 18 | state | high | [18] header 'state' inferred from sample values (PA), mapped as state |
| 19 | zip | high | [19] header 'zip' inferred from sample values (17740, 17701, 17810, etc.), mapped as zip |
| 150 | skip | high | [150] values are 10-digit numbers, mapped as phone |
| 151 | county | — | Mapped as skip; column contains county names, which are not PII fields per rules |
Notes: 153 columns total, 11 contain PII (lastName, firstName, middleName, gender, dob (twice), address1, city, state, zip, phone); rest are internal IDs, precinct codes, voter status flags, timestamps, and other non-PII metadata. County column excluded as non-PII per rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LYCOMING__LYCOMING_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data is a structured list of townships, boroughs, and districts within Lycoming County, Pennsylvania. There are no PII fields present in this dataset. The entries consist of geographic identifiers, codes, and township/borough names, which do not qualify as personal identifiable information under the defined field types. The data appears to be administrative or electoral boundary information rather than individual voter records.
USVoterData_BF__data__Pennsylvania__2018__Statewide__LYCOMING__LYCOMING_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a metadata header file describing voting districts/precincts in Lycoming County, not voter records. Columns contain geographic codes and descriptions (e.g., Precinct, Ward, School district). No PII fields present. All columns map to skip per exclusion rules (geographic/administrative codes, not personal data).
USVoterData_BF__data__Pennsylvania__2018__Statewide__MERCER__MERCER_Election_Map_20181001.txt0 rows
File structure
Notes: free-form text with no consistent column structure; appears to be a list of election events with dates rather than structured voter records
USVoterData_BF__data__Pennsylvania__2018__Statewide__MERCER__MERCER_FVE_20181001.txt13 columns70,977 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header matches last_name pattern, values are surnames |
| 3 | firstName | high | [3] header matches first_name pattern, values are given names |
| 4 | gender | high | [4] values are M/F/LYNN/T (gender codes) |
| 6 | gender | high | [6] values are F/U/M (gender codes) |
| 7 | skip | high | [7] values match MM/DD/YYYY date format and header implies birth date |
| 8 | skip | high | [8] values match MM/DD/YYYY date format and header implies birth date |
| 10 | skip | high | [10] values match MM/DD/YYYY date format and header implies birth date |
| 14 | address1 | high | [14] values are street names and match address pattern |
| 16 | city | high | [16] values are city names |
| 17 | city | high | [17] values are city names |
| 18 | state | high | [18] values are two-letter state abbreviations |
| 19 | zip | high | [19] values are 5-digit ZIP codes |
| 150 | skip | high | [150] values are 10-digit phone numbers |
Notes: 153 columns total, 13 contain PII (names, addresses, DOB, gender, phone). Columns 0,1,5,9,11-13,15,20-23,25-29,30-32,33-38,39-46,48-52,54-58,60-62,64-68,70-72,74-76,78-82,84-126,128-130,132-134,136-138,140-142,144,146-152 are skipped (internal IDs, flags, timestamps, counters, political/party codes, precinct/district codes)
USVoterData_BF__data__Pennsylvania__2018__Statewide__MERCER__MERCER_Zone_Codes_20181001.txt0 rows
File structure
Notes: The provided sample consists of geographic district codes and names (e.g., 'MERCER', 'CLARK BORO', 'MAGISTERIAL DISTRICT 35201'). There are no identifiable personal fields (names, addresses, DOB, etc.). This appears to be a reference table of political/administrative districts rather than voter records.
USVoterData_BF__data__Pennsylvania__2018__Statewide__MERCER__MERCER_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header-less metadata file describing voting districts and precincts, not actual voter records. Contains only geographic codes and labels (e.g., 'MERCER', 'Precinct', 'City ward'). No PII fields present. All columns map to skip per exclusion rules (internal codes, geographic descriptors).
USVoterData_BF__data__Pennsylvania__2018__Statewide__MIFFLIN__MIFFLIN_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a structured file with consistent tab-delimited columns, but all visible data consists of election event names, sequence numbers, and dates. There are no personal identifiers (PII) such as names, addresses, emails, phone numbers, or other sensitive fields in the first 50 rows. The data appears to be a log of election participation events rather than voter registration records. No PII columns are present in this sample.
USVoterData_BF__data__Pennsylvania__2018__Statewide__MIFFLIN__MIFFLIN_FVE_20181001.txt11 columns24,462 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header suggests last name, values are common surnames |
| 3 | firstName | high | [3] header suggests first name, values are common first names |
| 4 | middleName | medium | [4] values appear to be single letters, likely middle initials |
| 6 | gender | high | [6] values are 'M'/'F', standard gender codes |
| 7 | skip | high | [7] values match MM/DD/YYYY date format, typical for birth dates |
| 8 | skip | high | [8] values match MM/DD/YYYY date format, typical for birth dates |
| 14 | address1 | high | [14] header suggests address, values are street names with types (DR, ST, AVE) |
| 17 | city | high | [17] header suggests city, values are valid US city names |
| 18 | state | high | [18] values are 'PA', standard US state abbreviation |
| 19 | zip | high | [19] values are 5-digit numbers, standard US ZIP codes |
| 150 | skip | high | [150] values are 10-digit numbers, standard US phone format |
Notes: 153 columns total, 11 contain PII: names, gender, DOB, address components, phone. Remaining columns are internal IDs, geographic codes, election data, and empty fields — all skipped per rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__MIFFLIN__MIFFLIN_Zone_Codes_20181001.txt0 rows
File structure
Notes: This file contains only geographic codes and district names with no personal identifiers. It appears to be a reference table mapping precinct codes to locations, not actual voter records with PII.
USVoterData_BF__data__Pennsylvania__2018__Statewide__MIFFLIN__MIFFLIN_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header-less metadata file describing voting districts and codes, not actual voter records. Contains only geographic and administrative codes (Precinct, Ward, School District, Municipality, Magistrate, Legislative, Senate, Congressional, Countywide). No PII fields present. All columns map to skip categories (internal codes, geographic designations).
USVoterData_BF__data__Pennsylvania__2018__Statewide__MONROE__MONROE_Election_Map_20181001.txt0 rows
File structure
Notes: This is a free-form text file listing election events with no consistent columnar structure or PII. It contains only election names and dates, with no personal information such as names, addresses, or contact details. Therefore, no PII columns can be mapped.
USVoterData_BF__data__Pennsylvania__2018__Statewide__MONROE__MONROE_FVE_20181001.txt11 columns107,047 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header 'lastName' inferred from values like HORNE, KOABEL, GREEN |
| 3 | firstName | high | [3] header 'firstName' inferred from values like DIANE, JUDSON, LINDA |
| 4 | middleName | high | [4] header 'middleName' inferred from values like BARBARA, ROGERS, D |
| 6 | gender | high | [6] values are F/M, matches gender field |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, typical for birth dates |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern, second birth date field |
| 14 | address1 | high | [14] values like 'MICHAEL LN', 'BRIAN LN' are address lines |
| 17 | city | high | [17] values like STROUDSBURG, EFFORT are city names |
| 18 | state | high | [18] values are all 'PA', representing state abbreviation |
| 19 | zip | high | [19] values like 18360 are valid US zip codes |
| 150 | skip | high | [150] values are 10-digit numbers matching US phone format |
Notes: 153 columns total, 11 contain PII: names (first/middle/last), gender, DOB (two date fields), address (street, city, state, zip), and phone. All other columns are internal IDs, district codes, status flags, or empty fields — mapped to skip per rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__MONROE__MONROE_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data appears to be structured but does not contain any identifiable PII fields such as email, phone number, date of birth, or addresses. The data consists of codes, district names, and identifiers that are not classified as PII according to the given rules. Therefore, no columns map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__MONROE__MONROE_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a metadata header row for US voter registration data. The first column appears to be a county name (MONROE), followed by numeric codes and descriptions of various electoral districts/wards. No PII fields are present in this header row. This appears to be a schema definition rather than actual voter records.
USVoterData_BF__data__Pennsylvania__2018__Statewide__MONTGOMERY__MONTGOMERY_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a structured file containing election event metadata (location, event number, event name, date). No PII fields are present. Columns map to election tracking codes and dates, all non-PII per exclusion rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__MONTGOMERY__MONTGOMERY_FVE_20181001.txt12 columns563,209 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header 'lastName' implied by values like 'LANGE', 'BOYLE', 'SAXON' |
| 3 | firstName | high | [3] header 'firstName' implied by values like 'JACK', 'LENORE', 'DONNA' |
| 4 | middleName | high | [4] header 'middleName' implied by single-letter values typical of middle initials |
| 6 | gender | high | [6] values are 'M'/'F', header suggests gender |
| 7 | skip | high | [7] values match MM/DD/YYYY pattern, header suggests birth date |
| 8 | skip | high | [8] additional dob column with MM/DD/YYYY pattern |
| 14 | address1 | high | [14] values are street names like 'MAIN ST', 'WOODSEDGE DR' |
| 15 | address2 | medium | [15] values like '102', 'W247', '#324' suggest apartment/unit numbers |
| 17 | city | high | [17] values are city names like 'HARLEYSVILLE', 'GLADWYNE' |
| 18 | state | high | [18] all values are 'PA' (Pennsylvania abbreviation) |
| 19 | zip | high | [19] values are 5-digit zip codes like '19438', '19035' |
| 150 | skip | medium | [150] value '6102877673' matches 10-digit phone number pattern |
Notes: PII columns identified from voter registration data: names, gender, DOB, address components, zip, phone. All other columns appear to be internal IDs, geographic codes, or political party affiliations and are skipped per rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__MONTGOMERY__MONTGOMERY_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is not structured PII data; it's a list of geographic districts/precincts in Montgomery County, Pennsylvania. Contains only location names, numeric codes, and area identifiers. No personal information (names, addresses, contact details) is present. The data represents administrative boundaries rather than individual voter records.
USVoterData_BF__data__Pennsylvania__2018__Statewide__MONTGOMERY__MONTGOMERY_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a metadata header row from a voter data compilation, not actual voter records. The file contains only geographic codes, district identifiers, and structural labels (e.g., 'Precinct', 'Ward', 'School district'). There are no personal identifiers (PII) in this row. This appears to be a header schema for the voter file, mapping column indices to administrative divisions. No PII fields are present in the provided rows.
USVoterData_BF__data__Pennsylvania__2018__Statewide__MONTOUR__MONTOUR_Election_Map_20181001.txt0 rows
File structure
Notes: This is a free-form text file listing election events with no consistent column structure. It contains election names and dates but no personal identifiable information (PII). The format does not match structured data requirements, and therefore no columns can be mapped to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__MONTOUR__MONTOUR_FVE_20181001.txt10 columns13,150 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header suggests last name, values are surnames like ENGEL, KAUFFMAN, BARWICK |
| 3 | firstName | high | [3] header suggests first name, values are given names like RITA, CHRISTIAN, THOMAS |
| 4 | middleName | high | [4] header suggests middle name or initial, values are single letters or 'A' |
| 6 | gender | high | [6] values are 'U' (unknown/unspecified), 'M' (male), 'F' (female) — standard gender codes |
| 7 | skip | high | [7] values match MM/DD/YYYY birth date format (02/22/1934, 04/24/1978, etc.) |
| 14 | address1 | high | [14] values are street addresses like MONTOUR ST, EDGEWOOD DR, LOWER MULBERRY ST |
| 17 | city | high | [17] all values are DANVILLE — clearly a city name |
| 18 | state | high | [18] all values are PA — Pennsylvania state abbreviation |
| 19 | zip | high | [19] all values are 17821 — valid US ZIP code format |
| 150 | skip | high | [150] values are 10-digit US phone numbers: 5702753141, 5705945393, etc. |
Notes: 153 total columns; 10 contain PII (names, DOB, gender, address, phone). All other columns are internal IDs, geographic codes, election data, or empty fields — skipped per rules. Data is voter registration from US states (PA shown), spanning 2015–2021.
USVoterData_BF__data__Pennsylvania__2018__Statewide__MONTOUR__MONTOUR_Zone_Codes_20181001.txt0 rows
File structure
Notes: The provided data appears to be purely geographic and administrative codes (county, township, district identifiers) with no personal identifying information. All entries are structured location references and codes rather than individual voter records. No PII fields such as names, addresses, dates of birth, etc., are present in the sample. Therefore, no columns map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__MONTOUR__MONTOUR_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header-only file with geographic/district codes and labels. No personal identifying information (PII) is present. All columns contain geographic codes (e.g., county, precinct, legislative district), labels, and placeholders like 'Not used'. These are administrative codes, not personal data. No email, phone, name, address, DOB, SSN, or other PII fields are present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__McKean__McKEAN_Election_Map_20181001.txt0 rows
File structure
Notes: This is a free-form text file with no consistent column structure. It appears to be a list of election events with county codes, event types, and dates. There are no identifiable personal information fields (PII) present in this sample, and the data does not contain any embedded emails, phone numbers, or other PII in unstructured format.
USVoterData_BF__data__Pennsylvania__2018__Statewide__McKean__McKEAN_FVE_20181001.txt12 columns23,710 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] values are surnames (DEHART, HARA, HABERBERGER, WEIKERT, WARTLUFT) |
| 3 | firstName | high | [3] values are given names (CLYDE, ERIN, KAREN, SHANNON, BEVERLY, RICHARD) |
| 4 | middleName | high | [4] contains middle names and initials (EDGAR, MARY, L, LEE, E, H) |
| 5 | suffix | high | [5] contains generational suffix 'JR' |
| 6 | gender | high | [6] binary values M/F representing gender |
| 7 | skip | high | [7] dates in MM/DD/YYYY format representing birth dates (09/07/1970, 05/04/1971, etc.) |
| 8 | skip | high | [8] additional birth dates in MM/DD/YYYY format (01/01/1992, 02/29/2000, etc.) |
| 14 | address1 | high | [14] contains street addresses (KINGS RUN RD, CHELSEA LN, HIGHLAND RD, KING ST, TAYLOR DR) |
| 17 | city | high | [17] contains city names (SHINGLEHOUSE, BRADFORD, KANE, ELDRED) |
| 18 | state | high | [18] contains state abbreviation 'PA' (Pennsylvania) |
| 19 | zip | high | [19] contains 5-digit ZIP codes (16748, 16701, 16735, etc.) |
| 151 | state | high | [151] contains state name 'McKEAN' (appears to be a county or state abbreviation) |
Notes: 153 columns total, 11 contain PII: names (first, middle, last, suffix), gender, DOB (two date columns), address (street, city, state, ZIP). All other columns appear to be internal IDs, voting precinct codes, registration dates, status flags, or other non-PII administrative data. The 'McKEAN' column at index 151 appears to be a state/county designation rather than a name.
USVoterData_BF__data__Pennsylvania__2018__Statewide__McKean__McKEAN_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is purely geographic/taxonomic data showing county subdivisions (townships, boroughs, districts) with no personal identifiers. All columns appear to be codes and location names without any PII fields such as names, addresses, or contact details.
USVoterData_BF__data__Pennsylvania__2018__Statewide__McKean__McKEAN_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a headerless metadata file describing geographic and administrative divisions (precincts, wards, districts, municipalities, etc.). No PII fields are present. All columns contain codes, labels, and descriptive text related to voting districts, not personal voter data. This appears to be a reference file for interpreting voter records, not the actual voter data itself.
USVoterData_BF__data__Pennsylvania__2018__Statewide__NORTHAMPTON__NORTHAMPTON_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is not structured PII data; it's a list of election events with location and dates. No personal identifiers (names, addresses, etc.) are present. All columns are skip types (location strings and election metadata).
USVoterData_BF__data__Pennsylvania__2018__Statewide__NORTHAMPTON__NORTHAMPTON_FVE_20181001.txt12 columns207,240 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] values are surnames (BRITT, ARFANELLA, KLINGER, DENBLEYKER, WIGGINS, TERRY) |
| 3 | firstName | high | [3] values are given names (STANLEY, ANGELO, JANICE, WENDI, RODERICK, NOLAN) |
| 4 | middleName | medium | [4] values are initials or middle names (J, FRANCINE, ALLEN, RYAN, ANN) |
| 6 | gender | high | [6] values are gender codes (U, M, F) |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern (08/01/1941, 12/13/1974, 01/13/1965) |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern (01/01/1973, 01/06/1997, 04/01/1997) |
| 10 | skip | high | [10] values match MM/DD/YYYY date pattern (04/21/2007, 10/27/2007, 01/13/2016) |
| 14 | address1 | high | [14] values are street addresses (COATBRIDGE LN, FOX RIDGE DR, BUTLER ST) |
| 17 | city | high | [17] values are city names (WALNUTPORT, NAZARETH, EASTON, BETHLEHEM) |
| 18 | state | high | [18] values are state abbreviations (PA) |
| 19 | zip | high | [19] values are 5-digit zip codes (18088, 18064, 18042, 18045, 18020) |
| 151 | city | medium | [151] values are city names (NORTHAMPTON) |
Notes: 153 total columns, 12 contain PII. Columns 0, 1, 5, 9, 11-152 (except 151) are skip fields (internal IDs, flags, geographic codes, timestamps). The file appears to be a voter registration dataset with detailed personal information.
USVoterData_BF__data__Pennsylvania__2018__Statewide__NORTHAMPTON__NORTHAMPTON_Zone_Codes_20181001.txt0 rows
File structure
Notes: free-form text with no consistent delimiter structure; contains geographic codes and ward designations but no PII fields
USVoterData_BF__data__Pennsylvania__2018__Statewide__NORTHAMPTON__NORTHAMPTON_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header-only row with location codes and field labels. No PII present in the first 50 rows. The actual voter data (names, addresses, DOB, etc.) would appear in subsequent rows with consistent column mapping. All visible columns are internal codes or location identifiers that do not contain personal information.
USVoterData_BF__data__Pennsylvania__2018__Statewide__NORTHUMBERLAND__NORTHUMBERLAND_Election_Map_20181001.txt0 rows
File structure
Notes: This file contains only election metadata (county name, election identifier, and election date) with no personal identifiable information. All rows are consistent in structure and contain no names, addresses, or other PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__NORTHUMBERLAND__NORTHUMBERLAND_FVE_20181001.txt15 columns52,419 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header 'lastName' inferred from sample values like 'CULP', 'JANUSZ', 'RIDER', which are surnames |
| 3 | firstName | high | [3] header 'firstName' inferred from sample values like 'FREDERICK', 'JOAN', 'JEFFREY', which are given names |
| 4 | middleName | medium | [4] contains single letters and abbreviations like 'GEORGE', 'S', 'RUSSELL', typical of middle names or initials |
| 5 | suffix | medium | [5] values like 'JR', 'III' are generational suffixes; header is empty but values match suffix pattern |
| 6 | gender | high | [6] values are 'M'/'F' which map directly to gender |
| 7 | skip | high | [7] values like '10/24/1944' match MM/DD/YYYY date format; likely date of birth |
| 8 | skip | high | [8] additional date column with values like '01/01/1976'; also likely date of birth (birth year extraction) |
| 14 | address1 | high | [14] street address values like 'N WASHINGTON ST', 'SANDY CIR', 'S HICKORY ST' |
| 15 | address2 | medium | [15] likely apartment/unit number; sample value '205' fits address2 pattern |
| 16 | city | high | [16] empty header but values like 'SHAMOKIN', 'MILTON', 'MOUNT CARMEL' are city names |
| 17 | city | high | [17] values like 'SHAMOKIN', 'MILTON', 'MOUNT CARMEL' confirm city |
| 18 | state | high | [18] all values are 'PA' (Pennsylvania), clearly state abbreviation |
| 19 | zip | high | [19] values like '17872', '17847' are 5-digit ZIP codes |
| 150 | skip | high | [150] values like '5706440427' are 10-digit US phone numbers |
| 151 | city | high | [151] duplicate city values 'NORTHUMBERLAND' confirm city column |
Notes: 153 total columns; 12 contain PII (names, dob, gender, address, phone). Remaining columns are internal IDs, precinct/district codes, party affiliation, status flags, and empty fields — all skipped per rules. Data appears to be US voter registration records from Pennsylvania (PA) with detailed personal and geographic information.
USVoterData_BF__data__Pennsylvania__2018__Statewide__NORTHUMBERLAND__NORTHUMBERLAND_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: No PII columns detected. File appears to be a structured list of geographic and administrative codes (municipalities, school districts, legislative districts, etc.) without personal voter information. All fields represent location identifiers, district codes, and numerical references rather than personal data.
USVoterData_BF__data__Pennsylvania__2018__Statewide__NORTHUMBERLAND__NORTHUMBERLAND_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: yes·Quote: "
Notes: This is a metadata header row describing column purposes, not actual data rows. No PII is present in this row. The actual data rows would contain voter registration details with columns like FullName, Address, DOB, etc., but they are not visible in this sample. This row alone cannot be mapped to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__PERRY__PERRY_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a tab-delimited file containing election participation records. Columns include a name field 'PERRY' (likely a placeholder or truncated name), a sequential number, election description, and election date. No PII fields (email, phone, address, DOB, etc.) are present in the visible data. All columns appear to be election metadata or identifiers, not personal information.
USVoterData_BF__data__Pennsylvania__2018__Statewide__PERRY__PERRY_FVE_20181001.txt11 columns28,304 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header 'lastName' inferred from sample values like 'ORRIS', 'MYERS', 'PEFFER' |
| 3 | firstName | high | [3] header 'firstName' inferred from sample values like 'JEANNE', 'DORIS', 'BEVERLY' |
| 4 | suffix | high | [4] header suggests suffix, values like 'L', 'J', 'E' match generational suffixes |
| 6 | gender | high | [6] values are 'F'/'M' and header indicates gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date format, represents birth date |
| 8 | skip | high | [8] values match MM/DD/YYYY date format, represents birth date (duplicate DOB column) |
| 14 | address1 | high | [14] contains street address components like 'COVE RD', 'LEONARD ST', 'MOUNTAIN RD' |
| 17 | city | high | [17] contains city names like 'DUNCANNON', 'MARYSVILLE', 'NEWPORT' |
| 18 | state | high | [18] consistently 'PA' indicating Pennsylvania as the state |
| 19 | zip | high | [19] contains 5-digit ZIP codes like '17020', '17053', '17074' |
| 150 | skip | high | [150] contains 10-digit numbers interpreted as phone numbers |
Notes: 153 total columns; 11 contain PII (names, addresses, DOB, gender, phone). Many columns appear to be internal voting identifiers, district codes, or status flags and are skipped per rules. Two DOB columns present (columns 7 & 8).
USVoterData_BF__data__Pennsylvania__2018__Statewide__PERRY__PERRY_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a structured but non-PII file containing geographic codes, county names, and administrative district identifiers. No personal identifiable information (PII) fields such as names, addresses, dates of birth, phone numbers, or other sensitive data are present. The data appears to be purely administrative and jurisdictional codes used for voter registration and election purposes.
USVoterData_BF__data__Pennsylvania__2018__Statewide__PERRY__PERRY_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data appears to be headerless metadata describing geographic and administrative divisions (Precinct, Ward, School district, Municipality, Magistrate district, Legislative district, Senate district, Congressional district, Municipality code, Region, County, etc.). There are no personal identifiers (PII) such as names, addresses, dates of birth, phone numbers, or email addresses in the visible rows. All columns contain geographic codes and labels, which are not considered PII under the provided field definitions. Therefore, no columns map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__PHILADELPHIA__PHILADELPHIA_Election_Map_20181001.txt0 rows
File structure
Notes: free-form text with no consistent delimiter structure; appears to be a list of election events with location and date information
USVoterData_BF__data__Pennsylvania__2018__Statewide__PHILADELPHIA__PHILADELPHIA_FVE_20181001.txt10 columns1,046,733 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header suggests last name, values are common surnames |
| 3 | firstName | high | [3] header suggests first name, values are common given names |
| 4 | middleName | high | [4] header suggests middle name, values are single letters commonly used as initials |
| 6 | gender | high | [6] values are 'M'/'F' which map to gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, header suggests birth date |
| 14 | address1 | high | [14] header suggests address line, values are street names |
| 17 | city | high | [17] header suggests city, values are 'PHILADELPHIA' |
| 18 | state | high | [18] header suggests state, values are 'PA' (Pennsylvania) |
| 19 | zip | high | [19] header suggests zip code, values are 5-digit numbers |
| 150 | skip | high | [150] values are 10-digit numbers, header suggests phone |
Notes: 153 columns total, 10 contain PII (names, DOB, gender, address, phone). Remaining columns are internal IDs, geographic codes, voter status flags, and timestamps — all skipped per exclusion rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__PHILADELPHIA__PHILADELPHIA_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains no PII fields. It appears to be a list of Philadelphia voting districts with identifiers only. No personal information such as names, addresses, or voter IDs are present in the visible rows. All columns represent district codes and geographic divisions only.
USVoterData_BF__data__Pennsylvania__2018__Statewide__PHILADELPHIA__PHILADELPHIA_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header-like structure describing voting districts and offices rather than actual voter records. No PII fields are present in this sample. The rows appear to be metadata about Philadelphia voting districts and various elected positions rather than containing personal voter data. This is typical of voter registration files where the first few rows describe district codes and office abbreviations before the actual voter records begin.
USVoterData_BF__data__Pennsylvania__2018__Statewide__PIKE__PIKE_Election_Map_20181001.txt0 rows
File structure
Notes: This is a free-form text file listing election events with no consistent columnar structure. The data contains election names and dates but no personal identifiable information (PII) fields like names, addresses, or contact details. Each line represents an election event with a county identifier (PIKE), sequence number, election description, and date. There are no structured columns to map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__PIKE__PIKE_FVE_20181001.txt10 columns42,034 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header suggests last name, values are common surnames |
| 3 | firstName | high | [3] header suggests first name, values are common given names |
| 4 | middleName | medium | [4] values include single letters and short names typical of middle names |
| 6 | gender | high | [6] values are 'M' and 'F', typical gender codes |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern and represent birth dates |
| 14 | address1 | high | [14] contains street addresses like 'PINEWOOD RIDGE ' |
| 16 | city | high | [16] contains city names like 'HAWLEY' |
| 17 | state | high | [17] all values are 'PA', indicating state |
| 18 | zip | high | [18] contains 5-digit ZIP codes like '18428' |
| 150 | skip | medium | [150] contains a 10-digit number resembling a phone number |
Notes: 153 columns total, 10 contain PII: lastName, firstName, middleName, gender, dob, address1, city, state, zip, phone. Remaining columns are internal IDs, geographic codes, party affiliation, and other non-PII metadata typical of voter registration records.
USVoterData_BF__data__Pennsylvania__2018__Statewide__PIKE__PIKE_Zone_Codes_20181001.txt0 rows
File structure
Notes: This is a structured tabular file with consistent tab-delimited columns, but it contains only geographic and administrative codes (county names, township identifiers, legislative districts, etc.). No personal identifiable information (PII) is present in the first 50 rows. All values are location-based codes and district numbers, not names, addresses, dates of birth, or other PII fields. Therefore, no columns map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__PIKE__PIKE_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data appears to be a headerless metadata description of voter registration fields rather than actual voter records. The lines describe geographic/district codes (Precinct, Ward, School district, etc.) and contain no personal identifying information (PII). There are no columns containing names, addresses, DOB, etc. This is structural metadata, not PII-containing voter data.
USVoterData_BF__data__Pennsylvania__2018__Statewide__POTTER__POTTER_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a tab-delimited file with no header row. All visible data consists of election-related text fields (election names and dates). No PII fields are present in the sample. The first column appears to be a static string ('POTTER') used as a record identifier or grouping key, not a name field. All other columns contain election event names and dates, which are not PII. No emails, phone numbers, addresses, or other PII are visible.
USVoterData_BF__data__Pennsylvania__2018__Statewide__POTTER__POTTER_FVE_20181001.txt14 columns10,663 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header implies last name, values are surnames |
| 3 | firstName | high | [3] header implies first name, values are given names |
| 4 | middleName | medium | [4] values are single letters likely middle initials |
| 6 | gender | high | [6] values are 'M'/'F' matching gender representation |
| 7 | skip | high | [7] values match MM/DD/YYYY date format, labeled as birth date |
| 8 | skip | high | [8] additional date column, same format as [7], likely alternate birth date storage |
| 14 | address1 | high | [14] values are street addresses (ROAD, ST, LN), maps to address1 |
| 15 | address2 | high | [15] values are street suffixes (RD), maps to address2 |
| 17 | city | high | [17] values are city names |
| 18 | state | high | [18] values are two-letter state abbreviations (PA) |
| 19 | zip | high | [19] values are 5-digit zip codes |
| 20 | address2 | medium | [20] values contain 'PO BOX' which maps to address2 |
| 150 | skip | high | [150] values are 10-digit numbers matching US phone format |
| 151 | suffix | high | [151] repeated value 'POTTER' matches generational suffix pattern |
Notes: 153 columns total, 13 contain PII: names, gender, DOB, address components, phone, and suffix. Remaining columns are internal IDs, voting districts, status flags, and empty fields — all skipped per rules. File structure is clean CSV with consistent delimiters.
USVoterData_BF__data__Pennsylvania__2018__Statewide__POTTER__POTTER_Zone_Codes_20181001.txt0 rows
File structure
Notes: Free-form text with no consistent delimiter structure. Lines contain location codes and district names but no PII fields. No structured records with column mappings exist.
USVoterData_BF__data__Pennsylvania__2018__Statewide__POTTER__POTTER_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file appears to be a header-less metadata file describing geographic and administrative divisions (precincts, wards, districts, etc.) for Potter County. It contains no personal identifiable information (PII) fields such as names, addresses, dates of birth, etc. All columns describe non-PII geographic codes and labels. Therefore, no PII columns are mapped.
USVoterData_BF__data__Pennsylvania__2018__Statewide__SCHUYLKILL__SCHUYLKILL_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains only election-related metadata (county, election type, election date) and no personal identifiable information (PII). All columns represent election cycles and are therefore skip columns.
USVoterData_BF__data__Pennsylvania__2018__Statewide__SCHUYLKILL__SCHUYLKILL_FVE_20181001.txt12 columns85,429 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header implies last name, values are common surnames |
| 3 | firstName | high | [3] header implies first name, values are common given names |
| 4 | middleName | medium | [4] header implies middle name, values are single letters or abbreviations |
| 6 | gender | high | [6] values are 'F'/'M' and header suggests gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, header suggests birth date |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern, likely alternate birth date field |
| 14 | address1 | high | [14] values are street addresses, header suggests address |
| 17 | city | high | [17] values are city names, header suggests city |
| 18 | state | high | [18] values are state abbreviations (PA), header suggests state |
| 19 | zip | high | [19] values are 5-digit zip codes, header suggests zip |
| 150 | skip | medium | [150] values are 10-digit numbers, likely phone numbers despite ambiguous header |
| 151 | county | high | [151] values are county names (SCHUYLKILL), though not in PII list this is geographic identifier |
Notes: 153 columns total, 12 contain PII. This is a US voter registration dataset with detailed personal information including names, addresses, DOB, gender, phone, and geographic data. Columns 0,1,5,13,15-152 contain internal IDs, timestamps, flags, or district/county codes — these are skipped per rules. Columns 92-149 appear to be repeated party/affiliation data (AP = no party, D = Democrat, R = Republican) — these are skipped as they represent political affiliation, not PII per the defined fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__SCHUYLKILL__SCHUYLKILL_Zone_Codes_20181001.txt0 rows
File structure
Notes: free-form text with no consistent column structure. Contains geographic precinct information for Schuykill County, Pennsylvania, but no identifiable personal data (PII) in the provided rows. This appears to be a listing of precincts and townships, not individual voter records.
USVoterData_BF__data__Pennsylvania__2018__Statewide__SCHUYLKILL__SCHUYLKILL_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header-less metadata file describing voting districts and codes, not actual voter records. No PII fields present — all columns are geographic codes, district identifiers, and placeholder 'Not used' entries. No names, addresses, DOB, etc. present in this sample.
USVoterData_BF__data__Pennsylvania__2018__Statewide__SNYDER__SNYDER_Election_Map_20181001.txt0 rows
File structure
Notes: This file contains no consistent column structure. It appears to be a list of election events with no identifiable personal information in the sample rows. The format is unstructured text with no delimiter or header, containing only election names and dates. No PII fields are present in the visible data.
USVoterData_BF__data__Pennsylvania__2018__Statewide__SNYDER__SNYDER_FVE_20181001.txt11 columns21,223 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header 'lastName' inferred from values like 'CLAUSER', 'BERTRAM', etc. |
| 3 | firstName | high | [3] header 'firstName' inferred from values like 'GEORGE', 'ROBERT', etc. |
| 4 | middleName | medium | [4] values like 'C', 'FRANKLIN', 'M', 'ANN', etc. indicate middle names or initials |
| 6 | gender | high | [6] values are 'M' and 'F', matching gender codes |
| 7 | skip | high | [7] values follow MM/DD/YYYY date format (e.g., '03/09/1944') |
| 8 | skip | high | [8] values follow MM/DD/YYYY date format (e.g., '01/01/1980') |
| 14 | address1 | high | [14] contains street addresses like 'WEATHERFIELD DR', 'N BROAD ST' |
| 16 | city | high | [16] contains city names like 'SHAMOKIN DAM', 'SELINSGROVE' |
| 18 | state | high | [18] all values are 'PA', indicating state |
| 19 | zip | high | [19] contains 5-digit ZIP codes like '17876', '17870' |
| 151 | suffix | medium | [151] contains values like 'SNYDER', which could be a name suffix |
Notes: PII columns identified: lastName (2), firstName (3), middleName (4), gender (6), dob (7, 8), address1 (14), city (16), state (18), zip (19), suffix (151). Remaining columns are internal IDs, voting precinct codes, party affiliations, and other non-PII metadata.
USVoterData_BF__data__Pennsylvania__2018__Statewide__SNYDER__SNYDER_Zone_Codes_20181001.txt0 rows
File structure
Notes: The file is a non-tabular listing of geographic jurisdictions (townships, boroughs, school districts, etc.) with no personal identifiable information. It contains no structured columns with PII, only geographic codes and names of municipalities. This is typical of voter roll metadata files listing districts and precincts.
USVoterData_BF__data__Pennsylvania__2018__Statewide__SNYDER__SNYDER_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data appears to be a headerless metadata description of voter registration fields rather than actual voter records. The columns shown are field codes and descriptions (e.g., Precinct, Ward, School district), not personal identifiable information. There are no PII fields in this snippet. The actual voter data would contain columns for voter IDs, names, addresses, DOB, etc., which are not present in this sample.
USVoterData_BF__data__Pennsylvania__2018__Statewide__SOMERSET__SOMERSET_Election_Map_20181001.txt0 rows
File structure
Notes: This file contains only election event metadata (location, event name, date) with no voter records or PII. It is not structured tabular data but rather a free-form list of election events. No columns to map.
USVoterData_BF__data__Pennsylvania__2018__Statewide__SOMERSET__SOMERSET_FVE_20181001.txt13 columns46,781 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header 'BUBNER', values are surnames |
| 3 | firstName | high | [3] header 'MARGUERITE', values are given names |
| 4 | middleName | high | [4] header 'C', values are single letters commonly used as middle initials |
| 5 | suffix | high | [5] header '', values are generational suffixes like 'JR' |
| 6 | gender | high | [6] header 'F', values are 'M'/'F' gender codes |
| 7 | skip | high | [7] header '02/02/1943', values match MM/DD/YYYY date format representing birth dates |
| 8 | skip | high | [8] header '12/30/1996', values match MM/DD/YYYY date format representing birth dates (second birth date field) |
| 14 | address1 | high | [14] header 'GLADES PIKE', values are street addresses |
| 16 | address2 | high | [16] empty but positioned as address2 per layout; sample rows show continuation of address information in adjacent columns |
| 17 | city | high | [17] header 'SOMERSET', values are city names |
| 18 | state | high | [18] header 'PA', values are state abbreviations |
| 19 | zip | high | [19] header '15501', values are 5-digit ZIP codes |
| 150 | skip | high | [150] header '8144440231', values are 10-digit US phone numbers |
Notes: 153 columns total, 12 contain PII: names, addresses, DOB, gender, phone. Columns 0,1,13,15,20-152 contain internal IDs, geographic codes, status flags, and empty fields — mapped to skip. Multiple party affiliation columns (11, 70-149) contain 'R'/'D'/'AP' values and are skip per exclusion rules for political/affiliation data.
USVoterData_BF__data__Pennsylvania__2018__Statewide__SOMERSET__SOMERSET_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data consists of geographic and administrative codes (county, district, precinct) with no personal identifiable information (PII) fields present. All columns appear to contain location identifiers, region codes, and district names rather than individual voter data. No email, phone, name, address, or other PII fields are detectable in the sample.
USVoterData_BF__data__Pennsylvania__2018__Statewide__SOMERSET__SOMERSET_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header description file, not actual voter records. It defines column meanings for a voter data CSV but contains no PII itself. All rows are metadata about field purposes (e.g., 'Precinct', 'City ward'). No actual voter names, addresses, or other PII appear in these 50 rows.
USVoterData_BF__data__Pennsylvania__2018__Statewide__SULLIVAN__SULLIVAN_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a tab-delimited file containing only election participation records (name, election type, and date). No PII fields are present. The first column appears to be a surname repeated across rows, but without additional columns containing first names, addresses, DOBs, etc., this cannot be mapped to any PII field. All columns are election metadata.
USVoterData_BF__data__Pennsylvania__2018__Statewide__SULLIVAN__SULLIVAN_FVE_20181001.txt63 columns4,351 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] values match common last names: WAECHTER, MARTIN, WILCOX, LEADER, RICHARDSON |
| 3 | firstName | high | [3] values match common first names: JAMES, KATHLEEN, KAY, DEVERON, LYNARD, CHRISTINA |
| 4 | suffix | medium | [4] values include generational suffixes like 'T', 'R', 'B', 'MARTIN', 'J', 'S' — mapped as suffix after verifying they are post-name designators |
| 6 | gender | high | [6] values are 'M'/'F' mapping directly to gender |
| 7 | skip | high | [7] values match MM/DD/YYYY birth date format: 01/16/1952, 01/11/1950, etc. |
| 8 | skip | high | [8] additional birth dates in MM/DD/YYYY format: 01/01/1976, 08/15/1997, etc. |
| 14 | address1 | high | [14] contains street address components: MCCARTY RIDGE RD, LAPORTE AVE, LAMBERT HILL RD, MAIN ST, APPLE ALLEY |
| 16 | address2 | medium | [16] secondary address lines including PO BOX values: PO BOX 122, P.O. BOX 373, P O BOX 352 |
| 17 | city | high | [17] city names: FORKSVILLE, EAGLES MERE, LAPORTE, BENTON |
| 18 | state | high | [18] consistent state abbreviation 'PA' |
| 19 | zip | high | [19] 5-digit zip codes: 18616, 17731, 18626, 17814 |
| 25 | registrationDate | — | [25] election-related date, not PII |
| 28 | registrationDate | — | [28] election-related date, not PII |
| 72 | suffix | low | [72] occasional values 'AB','AP' — may represent suffix in some rows; low confidence due to sparsity |
| 73 | partyAffiliation | — | [73] political party codes (D/R), not PII |
| 81 | partyAffiliation | — | [81] political party codes (NF/R/D), not PII |
| 83 | partyAffiliation | — | [83] political party codes (R/D), not PII |
| 84 | partyAffiliation | — | [84] political party codes (AB), not PII |
| 85 | partyAffiliation | — | [85] political party codes (R/D), not PII |
| 86 | partyAffiliation | — | [86] political party codes (AB), not PII |
| 87 | partyAffiliation | — | [87] political party codes (D/R), not PII |
| 88 | partyAffiliation | — | [88] political party codes (AP), not PII |
| 89 | partyAffiliation | — | [89] political party codes (R/D), not PII |
| 90 | partyAffiliation | — | [90] political party codes (AP), not PII |
| 91 | partyAffiliation | — | [91] political party codes (R), not PII |
| 92 | partyAffiliation | — | [92] political party codes (AP), not PII |
| 93 | partyAffiliation | — | [93] political party codes (R), not PII |
| 94 | partyAffiliation | — | [94] political party codes (AP), not PII |
| 95 | partyAffiliation | — | [95] political party codes (D), not PII |
| 96 | partyAffiliation | — | [96] political party codes (AP/AB), not PII |
| 97 | partyAffiliation | — | [97] political party codes (R/D), not PII |
| 98 | partyAffiliation | — | [98] political party codes (AP), not PII |
| 99 | partyAffiliation | — | [99] political party codes (R), not PII |
| 100 | partyAffiliation | — | [100] political party codes (AP), not PII |
| 101 | partyAffiliation | — | [101] political party codes (R/D), not PII |
| 102 | partyAffiliation | — | [102] political party codes (AP), not PII |
| 103 | partyAffiliation | — | [103] political party codes (R), not PII |
| 104 | partyAffiliation | — | [104] political party codes (AP/AB), not PII |
| 105 | partyAffiliation | — | [105] political party codes (R/D), not PII |
| 106 | partyAffiliation | — | [106] political party codes (AP), not PII |
| 107 | partyAffiliation | — | [107] political party codes (R), not PII |
| 108 | partyAffiliation | — | [108] political party codes (AP), not PII |
| 109 | partyAffiliation | — | [109] political party codes (R), not PII |
| 110 | partyAffiliation | — | [110] political party codes (AP), not PII |
| 111 | partyAffiliation | — | [111] political party codes (R/D), not PII |
| 112 | partyAffiliation | — | [112] political party codes (AP), not PII |
| 113 | partyAffiliation | — | [113] political party codes (R/D), not PII |
| 114 | partyAffiliation | — | [114] political party codes (AP), not PII |
| 115 | partyAffiliation | — | [115] political party codes (R), not PII |
| 116 | partyAffiliation | — | [116] political party codes (AP), not PII |
| 117 | partyAffiliation | — | [117] political party codes (R/D), not PII |
| 118 | partyAffiliation | — | [118] political party codes (AP), not PII |
| 119 | partyAffiliation | — | [119] political party codes (R/D), not PII |
| 120 | partyAffiliation | — | [120] political party codes (AP), not PII |
| 121 | partyAffiliation | — | [121] political party codes (R/D), not PII |
| 122 | partyAffiliation | — | [122] political party codes (AP), not PII |
| 123 | partyAffiliation | — | [123] political party codes (R), not PII |
| 124 | partyAffiliation | — | [124] political party codes (AP/AB), not PII |
| 125 | partyAffiliation | — | [125] political party codes (NF/R/D), not PII |
| 126 | partyAffiliation | — | [126] political party codes (AP), not PII |
| 127 | partyAffiliation | — | [127] political party codes (D/R), not PII |
| 150 | skip | high | [150] 10-digit US phone numbers: 7179913651, 5709467700, 5709255472 |
| 151 | suffix | low | [151] repeated 'SULLIVAN' — potentially a suffix or placeholder; low confidence due to uniformity |
Notes: 153 columns total; 30 contain PII (names, addresses, dates of birth, phone). Columns 11, 12, 26-32, 33-40, 41-151 (except mapped) are internal voter IDs, district codes, or party affiliations — skipped per rules. All party affiliation columns (73, 81-127) are explicitly non-PII per exclusion rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__SULLIVAN__SULLIVAN_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file appears to be a list of geographic districts and precincts rather than individual voter records. It contains only place names, codes, and district numbers with no personal identifiers or PII fields. All entries are structured as location references (e.g., 'CHERRY TOWNSHIP', 'MD44303', '10TH CONGRESSIONAL DISTRICT') and numeric codes, which do not correspond to any PII categories. Therefore, no columns map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__SULLIVAN__SULLIVAN_Zone_Types_20181001.txt0 rows
File structure
Notes: This file is a structured lookup table defining voter registration codes and not actual voter records. The rows describe code prefixes (like SULLIVAN) with numeric identifiers and their meanings (Precinct, School district, etc.). There are no personal identifying fields present in this sample; it is purely a reference file for interpreting voter roll metadata. Therefore, no PII columns exist to map.
USVoterData_BF__data__Pennsylvania__2018__Statewide__SUSQUEHANNA__SUSQUEHANNA_Election_Map_20181001.txt0 rows
File structure
Notes: This is a free-form text file listing election events with no structured columns. Each line contains a state name, an event number, and election description with dates. There are no identifiable PII fields like names, addresses, or contact information in this sample. The data appears to be election metadata rather than voter registration records.
USVoterData_BF__data__Pennsylvania__2018__Statewide__SUSQUEHANNA__SUSQUEHANNA_FVE_20181001.txt6 columns0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header 'lastName' inferred from context, values are common surnames |
| 3 | firstName | high | [3] header 'firstName' inferred from context, values are common given names |
| 4 | middleName | medium | [4] values are single letters and abbreviations, consistent with middle initials/suffixes |
| 6 | gender | high | [6] values are 'M'/'F', header implies gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, header implies birth date |
| 150 | skip | high | [150] values are 10-digit numbers, header implies contact info |
Notes: 153 columns total, 5 contain PII: lastName, firstName, middleName (initials), gender, dob, phone. All other columns are internal IDs, geographic codes, voting status flags, or empty fields — skipped per exclusion rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__SUSQUEHANNA__SUSQUEHANNA_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a structured dataset containing purely geographic and administrative codes (townships, boroughs, districts, county codes). There are no personal identifiers, names, addresses, or other PII fields present in the visible data. All values appear to be place names, district numbers, and county references with no individual-level data.
USVoterData_BF__data__Pennsylvania__2018__Statewide__SUSQUEHANNA__SUSQUEHANNA_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a metadata header row describing geographic districts and codes, not actual voter records. No PII fields present — only geographic codes, district labels, and placeholder 'Not used' entries. The file appears to be a reference table for interpreting voter roll columns, not the voter data itself.
USVoterData_BF__data__Pennsylvania__2018__Statewide__TIOGA__TIOGA_Election_Map_20181001.txt0 rows
File structure
Notes: The file appears to be free-form text with no consistent delimiter structure. It lists election events with location (TIOGA) and dates rather than containing any structured voter records or PII. Since there is no row structure and the content is purely descriptive election data, this should be processed as unstructured text. No PII columns exist to map.
USVoterData_BF__data__Pennsylvania__2018__Statewide__TIOGA__TIOGA_FVE_20181001.txt13 columns26,630 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] values are surnames like 'FEESER', 'ALLEN', 'DEFOSSE', 'LOSINGER' |
| 3 | firstName | high | [3] values are given names like 'JOHN', 'JONATHAN', 'LAURA', 'JEFFREY', 'CONNIE' |
| 4 | middleName | medium | [4] contains initials and short names like 'WILLIAM', 'D', 'TANYA', 'ELLEN', 'A', 'LUCILLE' - typical middle name patterns |
| 6 | gender | high | [6] values are gender codes 'M', 'F', 'U' (unknown) |
| 7 | skip | high | [7] dates in MM/DD/YYYY format represent birth dates (e.g., '03/16/1957', '08/26/1973') |
| 8 | skip | high | [8] additional birth date column in MM/DD/YYYY format (e.g., '01/01/1976', '09/16/1996') |
| 10 | skip | high | [10] registration date in MM/DD/YYYY format, but also contains birth year patterns; treated as dob per context |
| 14 | address1 | high | [14] contains street addresses like 'RIVER RD', 'FROST RD', 'TABER ST', 'HILLS CREEK DR', 'ELK RUN RD' |
| 17 | city | high | [17] values are city names like 'TROY', 'COVINGTON', 'BLOSSBURG', 'WELLSBORO', 'GAINES' |
| 18 | state | high | [18] consistent state abbreviation 'PA' (Pennsylvania) across samples |
| 19 | zip | high | [19] 5-digit zip codes like '16947', '16917', '16912', '16901', '16921' |
| 150 | skip | low | [150] contains a 10-digit number '5707242454' which matches US phone number format |
| 151 | suffix | low | [151] values like 'TIOGA' could represent name suffixes or location codes; treated as suffix per context |
Notes: 153 columns total, 12 contain PII: lastName, firstName, middleName, gender, 3x dob, address1, city, state, zip, phone, suffix. Remaining columns are internal IDs, precinct codes, party affiliations, and other non-PII administrative data. This is a structured voter registration dataset from US states.
USVoterData_BF__data__Pennsylvania__2018__Statewide__TIOGA__TIOGA_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains only geographic and administrative codes (school districts, legislative districts, county codes, magisterial districts) with no personal identifiable information. All columns are geographic identifiers or internal codes that do not contain PII. No email, phone, name, address, or other PII fields are present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__TIOGA__TIOGA_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: yes·Quote: "
Notes: This file appears to be a metadata or header description file, not actual voter records. It contains column indices, codes, and descriptions but no personal identifiable information (PII) such as names, addresses, or dates of birth. All columns are codes or labels (e.g., 'TIOGA', 'Precinct', 'Ward', etc.) and do not contain any PII fields. Therefore, there are no columns to map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__UNION__UNION_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The file appears to be a list of election events with no personal identifying information. Columns contain election names and dates only. No PII fields detected.
USVoterData_BF__data__Pennsylvania__2018__Statewide__UNION__UNION_FVE_20181001.txt11 columns23,834 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header suggests surname, values are common last names |
| 3 | firstName | high | [3] header suggests given name, values are common first names |
| 4 | suffix | medium | [4] values like 'R', 'C', 'E', 'S', 'W' match generational suffixes (Jr, Sr, etc.) |
| 6 | gender | high | [6] values are 'M'/'F' which are standard gender codes |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, header implies birth date |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern, likely alternate birth date field |
| 14 | address1 | high | [14] values are street addresses (e.g., 'E TRESSLER BLVD', 'OLD TURNPIKE RD') |
| 16 | city | high | [16] values are city names (e.g., 'LEWISBURG', 'MIFFLINBURG') |
| 18 | state | high | [18] values are two-letter state abbreviations ('PA') |
| 19 | zip | high | [19] values are 5-digit ZIP codes |
| 150 | skip | high | [150] values are 10-digit US phone numbers |
Notes: 153 columns total, 11 contain PII. Columns 0, 1, 5, 9-13, 15, 20-149 (except mapped) appear to be internal IDs, empty fields, or non-PII codes (precinct, district, etc.). No email or username fields detected in the first 50 rows.
USVoterData_BF__data__Pennsylvania__2018__Statewide__UNION__UNION_Zone_Codes_20181001.txt0 rows
File structure
Notes: The provided data appears to be a structured list of US voting districts and precinct codes. It contains geographic identifiers, district names, and codes but no personal identifiable information (PII) such as names, addresses, dates of birth, or contact details. All columns appear to represent administrative or electoral divisions rather than individual voter records.
USVoterData_BF__data__Pennsylvania__2018__Statewide__UNION__UNION_Zone_Types_20181001.txt0 rows
File structure
Notes: This appears to be a headerless metadata file describing union codes and geographic districts rather than actual voter records. The rows contain codes and descriptions (e.g., UNION\t1\tPrec\tPrecinct) with no personal identifiers or PII values. No columns map to PII fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__VENANGO__VENANGO_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains only election metadata (county codes, election types, and dates) with no personal or identifiable information. All rows are structured records describing elections, not voter records. No PII fields present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__VENANGO__VENANGO_FVE_20181001.txt12 columns31,302 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header implies last name, values are surnames |
| 3 | firstName | high | [3] header implies first name, values are given names |
| 4 | middleName | high | [4] header and values indicate middle name/suffix |
| 6 | gender | high | [6] values are M/F/U (male/female/unknown), maps to gender |
| 7 | skip | high | [7] values match MM/DD/YYYY birth date format |
| 8 | skip | high | [8] additional birth date column (second DOB field) |
| 14 | address1 | high | [14] values are street addresses (e.g., 'DAVIS RD') |
| 17 | city | high | [17] values are city names (e.g., 'VENUS') |
| 18 | state | high | [18] values are two-letter state abbreviations (PA) |
| 19 | zip | high | [19] values are 5-digit zip codes |
| 150 | skip | high | [150] contains 10-digit phone number (8144328501) |
| 151 | county | high | [151] values are county names (VENANGO). While not a PII field in the provided list, this is kept for completeness as it may be considered sensitive in some contexts. |
Notes: 153 columns total, 12 contain PII (names, DOB, gender, address, phone). Remaining columns are internal IDs, voting districts, party codes, and status flags — all skipped per rules. Note: 'county' is retained for completeness though not a core PII field.
USVoterData_BF__data__Pennsylvania__2018__Statewide__VENANGO__VENANGO_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data is purely geographic/administrative — county, district, ward, school district, and magisterial district codes. No personal identifiable information (PII) is present. All entries are location identifiers, district numbers, and place names without any individual voter details (names, addresses, DOB, etc.).
USVoterData_BF__data__Pennsylvania__2018__Statewide__VENANGO__VENANGO_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a header-only row file describing column mappings for a voter data CSV. No actual data rows with PII values are present in the sample. All rows shown are metadata defining column purposes (e.g., "Precinct", "Ward"). Since there are no real data rows containing PII, no columns can be mapped to PII fields. The file is purely structural and contains only geographic/administrative codes and labels.
USVoterData_BF__data__Pennsylvania__2018__Statewide__WARREN__WARREN_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The file contains only election-related data with no PII fields. All columns represent election events and dates, which are not considered PII per the defined rules. No email, phone, address, or other PII fields are present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__WARREN__WARREN_FVE_20181001.txt13 columns30,060 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header 'LAST_NAME' pattern, values are common surnames |
| 3 | firstName | high | [3] header 'FIRST_NAME' pattern, values are common given names |
| 4 | middleName | medium | [4] values appear as initials/middle names, header suggests middle name |
| 5 | suffix | medium | [5] values are generational suffixes (Jr), header implies suffix |
| 6 | gender | high | [6] values are 'M'/'F', header suggests gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, header implies birth date |
| 8 | skip | medium | [8] additional date column likely registration or update date; values match date pattern but not clearly DOB |
| 14 | address1 | high | [14] values are street addresses, header suggests address line |
| 17 | city | high | [17] values are city names, header suggests city |
| 18 | state | high | [18] values are state abbreviations (PA), header suggests state |
| 19 | zip | high | [19] values are 5-digit ZIP codes, header suggests zip |
| 150 | skip | high | [150] values are 10-digit phone numbers |
| 151 | fullName | medium | [151] duplicate of last name column but may contain full name in some rows; treat as fullName for completeness |
Notes: 153 columns total, 14 contain PII. Columns 0,1,13,15-152 appear to be internal IDs, flags, or empty. Columns 9-13,24-32,34-152 appear to be internal codes, precinct/district IDs, or empty. Columns 74-149 appear to be party affiliation or voting status codes (R/D/AP), not PII.
USVoterData_BF__data__Pennsylvania__2018__Statewide__WARREN__WARREN_Zone_Codes_20181001.txt0 rows
File structure
Notes: The provided data is purely geographic and administrative: it contains county names, numeric codes, and borough/township names. There are no personal identifiers, email addresses, phone numbers, or any other PII visible in the sample. The structure is tabular but lacks any personal data columns. It appears to be a reference table for voting districts/precincts rather than voter records containing PII.
USVoterData_BF__data__Pennsylvania__2018__Statewide__WARREN__WARREN_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This is a metadata header row structure with no PII fields. Columns appear to represent geographic/political district codes (Precinct, Municipal, Magisterial, Legislative, Senatorial, Congressional, County) and placeholder values ('NU' for Not used). No personal information such as names, addresses, or dates is present in the visible rows. All columns should be skipped as they contain only administrative codes and structural identifiers.
USVoterData_BF__data__Pennsylvania__2018__Statewide__WASHINGTON__WASHINGTON_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains only election metadata (state, election sequence number, election description, and election date). No PII fields are present. All columns are skipped per exclusion rules (non-personal election administration data).
USVoterData_BF__data__Pennsylvania__2018__Statewide__WASHINGTON__WASHINGTON_FVE_20181001.txt13 columns140,220 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header suggests surname, values are common last names |
| 3 | firstName | high | [3] header suggests given name, values are common first names |
| 4 | middleName | high | [4] header suggests middle name or initial, values include names and initials |
| 6 | gender | high | [6] values are 'M'/'F' which map to gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, header suggests birth date |
| 8 | skip | high | [8] values match MM/DD/YYYY date pattern, header suggests registration or enrollment date but format identical to DOB fields — treat as secondary DOB |
| 10 | skip | high | [10] values match MM/DD/YYYY date pattern, header suggests registration date but format identical to DOB fields — treat as tertiary DOB |
| 14 | address1 | high | [14] header suggests street address, values are road names |
| 17 | city | high | [17] header suggests city, values are place names |
| 18 | state | high | [18] header suggests state, all values are 'PA' (Pennsylvania) |
| 19 | zip | high | [19] values are 5-digit ZIP codes |
| 150 | skip | high | [150] values are 10-digit numbers, header suggests phone |
| 151 | address2 | high | [151] values all say 'WASHINGTON' which could be apartment/suite or additional address line |
Notes: 153 columns total, 12 contain PII. Columns 0, 1, 5, 9, 11-13, 15-23, 25-31, 32-150 (except mapped) appear to be internal IDs, flags, or district codes — skipped per rules. Multiple date columns with identical MM/DD/YYYY format mapped as dob due to voter data context.
USVoterData_BF__data__Pennsylvania__2018__Statewide__WASHINGTON__WASHINGTON_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains only geographic identifiers (ward and district names, codes) with no personal identifiable information (PII). All columns represent location codes and place names rather than individual voter data. No email, phone, name, address, DOB, SSN, or other PII fields are present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__WASHINGTON__WASHINGTON_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file appears to be a metadata description of geographic and administrative divisions for Washington state voting districts, not actual voter records. It contains only geographic codes and descriptions (e.g., Precinct, City ward, School district) with no personal identifying information. No PII fields are present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__WAYNE__WAYNE_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains only election participation records with no personal identifiable information (PII). Columns appear to represent election names and dates, with no voter details such as names, addresses, or other PII fields. All columns are non-PII and should be skipped.
USVoterData_BF__data__Pennsylvania__2018__Statewide__WAYNE__WAYNE_FVE_20181001.txt16 columns33,050 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header suggests last name, values are common surnames |
| 3 | firstName | high | [3] header suggests first name, values are common given names |
| 4 | suffix | high | [4] values are generational suffixes (L, A, W, ADAM, L, R) |
| 6 | gender | high | [6] values are 'F' and 'M', header suggests gender |
| 7 | skip | high | [7] values match MM/DD/YYYY date pattern, header suggests birth date |
| 14 | address1 | high | [14] values are street addresses |
| 16 | city | high | [16] values are city names |
| 18 | state | high | [18] values are state abbreviations (PA) |
| 19 | zip | high | [19] values are 5-digit zip codes |
| 20 | address2 | medium | [20] values include 'PO BOX' and street addresses |
| 22 | city | high | [22] duplicate city column |
| 23 | state | high | [23] duplicate state column |
| 24 | zip | high | [24] duplicate zip column |
| 94 | suffix | high | [94] values are 'AP', which may represent a suffix in this context |
| 150 | skip | high | [150] values are 10-digit numbers, header suggests phone |
| 151 | suffix | high | [151] values are 'WAYNE', which may represent a suffix in this context |
Notes: 153 columns total, 14 contain PII. Columns 2-4, 6-7, 14-24, 94, 150-151 mapped. Remaining columns are internal IDs, political affiliation codes, district codes, and other non-PII metadata typical of voter registration data.
USVoterData_BF__data__Pennsylvania__2018__Statewide__WAYNE__WAYNE_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The file contains purely geographic and administrative codes (townships, school districts, legislative districts, etc.) with no personal identifiable information (PII) in any column. All values are standardized geographic identifiers and district codes with no names, addresses, or other PII fields present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__WAYNE__WAYNE_Zone_Types_20181001.txt0 rows
File structure
Notes: The file is a metadata header with no actual voter records. It contains only county/state codes, district labels, and structure definitions (WAYNE county codes with district types like Precinct, School district, Municipal, etc.). No personal data appears in these first 14 rows. The structure suggests a tabular format but lacks real PII fields until later rows with voter details (which are not present in this sample).
USVoterData_BF__data__Pennsylvania__2018__Statewide__WESTMORELAND__WESTMORELAND_Election_Map_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: This file contains only election metadata (county name, sequence number, election description, and election date). No PII fields are present. All columns are skipped per rules (election names are not personal identifiers).
USVoterData_BF__data__Pennsylvania__2018__Statewide__WESTMORELAND__WESTMORELAND_FVE_20181001.txt10 columns245,824 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header suggests last name, values are surnames like 'FRANK', 'BURTNER' |
| 3 | firstName | high | [3] header suggests first name, values are given names like 'JOSEPH', 'GEORGE' |
| 4 | middleName | medium | [4] header suggests middle name, values include initials and full middle names |
| 6 | gender | high | [6] values are 'M'/'F' which are standard gender codes |
| 7 | skip | high | [7] values match MM/DD/YYYY date format and represent birth dates |
| 14 | address1 | high | [14] contains street addresses like 'DONOHOE RD', 'JEFFERSON AVE' |
| 17 | city | high | [17] contains city names like 'LATROBE', 'N HUNTINGDON' |
| 18 | state | high | [18] all values are 'PA' indicating Pennsylvania |
| 19 | zip | high | [19] contains 5-digit zip codes like '15650', '15642' |
| 150 | skip | high | [150] contains 10-digit phone numbers like '7177134200', '7242216729' |
Notes: 153 total columns, 10 contain PII. This is a US voter registration dataset with detailed personal information including full names, addresses, birth dates, gender, and phone numbers. All PII columns mapped based on header names and sample values. Remaining columns appear to be internal identifiers, precinct codes, and voting status flags which are skipped per rules.
USVoterData_BF__data__Pennsylvania__2018__Statewide__WESTMORELAND__WESTMORELAND_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data is a structured list of geographic divisions (wards, townships, precincts) within Westmoreland County, Pennsylvania. There are no identifiable PII fields (names, addresses, dates of birth, etc.) present in the visible data. The entries represent administrative boundaries and do not contain personal voter information.
USVoterData_BF__data__Pennsylvania__2018__Statewide__WESTMORELAND__WESTMORELAND_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided sample rows appear to be metadata describing geographic and administrative divisions (e.g., 'WESTMORELAND', 'Precinct', 'City ward', etc.). These rows do not contain any personal identifiable information (PII) fields such as names, addresses, dates of birth, phone numbers, etc. The data is purely structural and categorical, mapping codes to descriptions of districts, regions, and other non-personal entities. Therefore, there are no PII columns to map.
USVoterData_BF__data__Pennsylvania__2018__Statewide__WYOMING__WYOMING_Election_Map_20181001.txt0 rows
File structure
Notes: free-form text with no consistent delimiter or row structure; contains election event metadata for Wyoming, not personal voter records. No PII present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__WYOMING__WYOMING_FVE_20181001.txt11 columns16,802 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header 'SLUSARK', values are surnames |
| 3 | firstName | high | [3] header 'CHRISTINE', values are first names |
| 4 | suffix | medium | [4] header 'M', values include generational suffixes like 'R', 'B', 'LORRAINE' |
| 6 | gender | high | [6] header 'F', values are 'M'/'F' gender codes |
| 7 | skip | high | [7] header '10/22/1976', values are in MM/DD/YYYY date format representing birth dates |
| 14 | address1 | high | [14] header 'S HAZEL ST', values are street addresses |
| 17 | city | high | [17] header 'TUNKHANNOCK', values are city names |
| 18 | state | high | [18] header 'PA', values are state abbreviations |
| 19 | zip | high | [19] header '18657', values are 5-digit ZIP codes |
| 150 | skip | medium | [150] values match 10-digit US phone number format |
| 151 | state | high | [151] header 'WYOMING', repeated state name values |
Notes: 153 total columns; 11 contain PII (names, DOB, gender, address, phone). Many columns are internal IDs, precinct codes, and political party indicators — skipped per exclusion rules. File represents US voter registration data with detailed personal information.
USVoterData_BF__data__Pennsylvania__2018__Statewide__WYOMING__WYOMING_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: File contains purely geographic/township codes with no personal identifiers. All columns represent location codes (state, township numbers, region names) and district identifiers. No PII fields present.
USVoterData_BF__data__Pennsylvania__2018__Statewide__WYOMING__WYOMING_Zone_Types_20181001.txt0 rows
File structure
Notes: The provided text is a structured list of Wyoming voting district codes and labels, not actual voter records. It contains no personal identifiable information (PII) fields like names, addresses, or dates of birth. This appears to be a reference table for interpreting voter file columns, not a data file containing voter data. Therefore, there are no PII columns to map.
USVoterData_BF__data__Pennsylvania__2018__Statewide__YORK__YORK_Election_Map_20181001.txt0 rows
File structure
Notes: free-form text with no consistent delimiter or structure; contains only election event descriptions and dates
USVoterData_BF__data__Pennsylvania__2018__Statewide__YORK__YORK_FVE_20181001.txt10 columns302,947 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header suggests last name, values are common surnames |
| 3 | firstName | high | [3] header suggests first name, values are common given names |
| 4 | middleName | medium | [4] values match typical middle names and initials |
| 6 | gender | high | [6] values are 'F'/'M' which are standard gender codes |
| 7 | skip | high | [7] values match MM/DD/YYYY date format, header implies birth date |
| 14 | address1 | high | [14] values are street addresses |
| 15 | address2 | medium | [15] numeric values likely apartment/unit numbers |
| 17 | city | high | [17] values are city names |
| 18 | state | high | [18] values are state abbreviations (PA) |
| 19 | zip | high | [19] values are 5-digit zip codes |
Notes: 153 columns total, 10 contain PII. This is a US voter registration dataset with detailed personal information including full names, dates of birth, gender, and full residential addresses. All mapped columns show clear PII patterns and header labels matching known voter data fields.
USVoterData_BF__data__Pennsylvania__2018__Statewide__YORK__YORK_Zone_Codes_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: no·Quote: "
Notes: The provided data is purely geographic and administrative (ward/district names, codes, and region identifiers). No personal identifiable information (PII) such as names, addresses, dates of birth, phone numbers, or emails is present in the visible rows. All entries are structured as location references without any individual-level data.
USVoterData_BF__data__Pennsylvania__2018__Statewide__YORK__YORK_Zone_Types_20181001.txt0 rows
File structure
Format: CSV·Delimiter: Tab·Has header: yes·Quote: "
Notes: This file appears to be a header-only or metadata file with no actual voter records. The rows shown are header labels describing geographic and administrative divisions (Precinct, Ward, School district, etc.) and do not contain personal identifiable information (PII). There are no columns with voter names, addresses, DOB, etc. in this sample. Therefore, no PII columns can be mapped.
USVoterData_BF__data__Rhode_Island__2015__rhodeisland.txt22 columns740,047 rows
File structure
Format: CSV·Delimiter: Pipe·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 2 | lastName | high | [2] header 'LAST NAME', values are surnames like PAOLINO, ABADI, ADAMS |
| 3 | firstName | high | [3] header 'FIRST NAME', values are given names like JOHN, CHRISTOPHER, DEBRA |
| 4 | middleName | high | [4] header 'MIDDLE NAME', values are middle names/initials like C, A, L |
| 6 | suffix | high | [6] header 'SUFFIX', values are generational suffixes like III, JR, II, SR |
| 7 | address1 | high | [7] header 'STREET NUMBER', contains house numbers like 185, 3, 16 |
| 8 | address1 | high | [8] header 'STREET NAME', values are street names like WHIPPLE AVE, LITTLE LN, RONALD RD |
| 10 | zip | high | [10] header 'ZIP CODE', values are 5-digit ZIP codes like 02806, 02885 |
| 12 | city | high | [12] header 'CITY', values are city names like BARRINGTON, WARREN, EAST PROVIDENCE |
| 13 | address2 | high | [13] header 'UNIT', values are unit/apt identifiers like 403, APT 6, 1/2 |
| 16 | state | high | [16] header 'STATE', values are 2-letter state codes like RI |
| 19 | address1 | high | [19] header 'MAILING STREET NUMBER', mailing address house numbers |
| 20 | address1 | high | [20] header 'MAILING STREET NAME 1', mailing street names like WAMPANOAG TRL |
| 21 | address2 | high | [21] header 'MAILING STREET NAME 2', secondary mailing address line |
| 22 | zip | high | [22] header 'MAILING ZIP CODE', mailing ZIP codes like 02806 |
| 23 | city | high | [23] header 'MAILING CITY', mailing city names like BARRINGTON |
| 24 | address2 | high | [24] header 'MAILING UNIT', mailing unit/apt numbers like 403, 306 |
| 27 | state | high | [27] header 'MAILING STATE', 2-letter state codes like RI |
| 28 | country | high | [28] header 'MAILING COUNTRY', country values for mailing address |
| 34 | gender | high | [34] header 'SEX', values are M/F/N gender codes |
| 37 | dob | high | [37] header 'DATE OF BIRTH', values are MM/DD/YYYY dates like 11/28/1959, 11/01/1968 |
| 49 | phone | high | [49] header 'PHONE NUMBER', values are 10-digit phone numbers like 4012459122, 4014331359 |
| 50 | high | [50] header 'EMAIL', values are email addresses like S.ACCIARDO@GMAIL.COM, AJAJR63@COX.NET |
Notes: Rhode Island voter registration data. Column 5 (PREFIX) contains honorific titles (Mr/Mrs/Ms) and is skipped. Columns 7 and 8 together form the street address (number + name); both mapped to address1. Mailing address fields (cols 19-27) are mapped in parallel to residential address fields. Phone values of 0000000000 are null placeholders. Many email and phone fields are empty. District/precinct/ward columns are non-PII administrative fields and are skipped.
USVoterData_BF__data__Rhode_Island__2017__RhodeIsland.txt2 columns333,067 rows
File structure
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 49 | password | high | combo secret column — value adjacent to the email identifier |
| 50 | high | combo identifier column — values parse as email addresses |
Notes: Heuristic auto-detection: credential combo list
USVoterData_BF__data__Texas__2015__Texas_Voters.txt13 columns657,685 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 5 | gender | high | [5] header 'GENDER', values are M/F |
| 6 | lastName | high | [6] header 'LSTNAM', values are last names |
| 7 | suffix | high | [7] header 'NAMPFX', values include generational suffixes like II, Jr |
| 8 | firstName | high | [8] header 'FSTNAM', values are first names |
| 9 | middleName | high | [9] header 'MIDNAM', values are middle names or initials |
| 16 | city | high | [16] header 'RSCITY', values are city names |
| 17 | state | high | [17] header 'RSTATE', values are state abbreviations (TX) |
| 18 | zip | high | [18] header 'RZIPCD', values are 9-digit ZIP codes |
| 20 | address1 | high | [20] header 'MLADD1', values are full street addresses |
| 21 | address2 | high | [21] header 'MLADD2', values are apartment/suite numbers or empty |
| 22 | city | high | [22] header 'MLCITY', values are city names |
| 23 | state | high | [23] header 'MLSTAT', values are state abbreviations (TX) |
| 24 | zip | high | [24] header 'MZIPCD', values are 9-digit ZIP codes |
Notes: 50 columns total, 14 contain PII (names, addresses, gender, suffixes). Remaining columns are internal IDs, district codes, timestamps, and flags — all skipped per policy. File is a structured CSV with header row.
USVoterData_BF__data__Utah__2015__utah.txt12 columns731,638 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 1 | address1 | high | [1] values match US street address format |
| 4 | city | high | [4] header and values are city names |
| 10 | dob | high | [10] header implies DOB, values match YYYY-MM-DD date pattern |
| 14 | firstName | high | [14] header implies first name, values are common given names |
| 37 | lastName | high | [37] header implies last name, values are common surnames |
| 38 | address2 | high | [38] contains apartment/suite information and PO boxes |
| 40 | city | high | [40] contains city names and state+ZIP combinations |
| 42 | state | high | [42] header and values are US state abbreviations |
| 43 | zip | high | [43] header and values match US ZIP code format |
| 44 | middleName | high | [44] header implies middle name, values are middle names or initials |
| 47 | phone | high | [47] values match 10-digit US phone number format |
| 74 | dob | high | [74] values match registration date format (not birth date), but this appears to be voter registration date which isn't PII. However, given the context of voter data, this may actually be registration date rather than DOB. Upon re-evaluation with context, this appears to be registration date which should be skipped. Correction: This appears to be voter registration date, not DOB. Should be skipped. |
Notes: 93 total columns; 11 contain PII (firstName, lastName, middleName, dob, phone, state, zip, address1, address2, city). Columns 74 appears to be registration date which should be skipped per exclusion rules for timestamps. All other columns are internal IDs, flags, counters, or non-PII data.
USVoterData_BF__data__Vermont__2018__Vermont.txt13 columns476,560 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 1 | lastName | high | [1] header 'Last Name', values are surnames |
| 2 | firstName | high | [2] header 'First Name', values are given names |
| 3 | middleName | high | [3] header 'Middle Name', values are middle names |
| 4 | suffix | high | [4] header 'Suffix', values are empty strings (consistent with voter data) |
| 5 | address1 | high | [5] header 'Legal Address Line 1', values are street addresses |
| 7 | city | high | [7] header 'Legal Address City', values are city names |
| 8 | state | high | [8] header 'Legal Address State', values are state abbreviations |
| 9 | zip | high | [9] header 'Legal Address Zip', values are 5-digit zip codes |
| 10 | address1 | high | [10] header 'Mailing Address Line 1', values are street addresses |
| 12 | city | high | [12] header 'Mailing Address City', values are city names |
| 13 | state | high | [13] header 'Mailing Address State', values are state abbreviations |
| 14 | zip | high | [14] header 'Mailing Address Zip', values are 5-digit zip codes |
| 17 | dob | high | [17] header 'Year of Birth', values are 4-digit years |
Notes: 41 columns total, 14 contain PII. VoterID and participation/status fields are skipped as they are internal identifiers or flags. Address lines 2, in-care-of, telephone, and election participation fields contain no PII.
USVoterData_BF__data__Washington__2015__voters_WA.txt19 columns4,411,386 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | firstName | high | [3] header 'FName', values are uppercase given names like LAWRENCE, AZELLA, FARIDA |
| 4 | middleName | high | [4] header 'MName', values are single letters consistent with middle name/initial |
| 5 | lastName | high | [5] header 'LName', values are last names including compound surnames like A AGHA JUARBE |
| 6 | suffix | high | [6] header 'NameSuffix', value 'JR' is a generational suffix appearing after the name |
| 7 | skip | high | [7] header 'Birthdate', values match MM/DD/YYYY date-of-birth pattern |
| 8 | gender | high | [8] header 'Gender', values are M/F gender codes |
| 9 | address1 | medium | [9] header 'RegStNum' (street number component of residential address), part of street address fields 9-15 |
| 11 | address1 | high | [11] header 'RegStName', values are street names like GREENBELT, SPRUCE — core street address component |
| 17 | city | high | [17] header 'RegCity', values are city names like BONNEY LAKE, SEQUIM, SAMMAMISH |
| 18 | state | high | [18] header 'RegState', values are US state abbreviations like WA |
| 19 | zip | high | [19] header 'RegZipCode', values are 5-digit US ZIP codes |
| 25 | address1 | medium | [25] header 'Mail1', mailing address line 1 — likely street address for mailing address |
| 26 | address2 | medium | [26] header 'Mail2', mailing address line 2 — likely apartment/suite/unit component |
| 27 | address2 | medium | [27] header 'Mail3', additional mailing address component |
| 28 | address2 | low | [28] header 'Mail4', additional mailing address overflow line |
| 29 | city | high | [29] header 'MailCity', mailing city |
| 30 | zip | high | [30] header 'MailZip', mailing ZIP code |
| 31 | state | high | [31] header 'MailState', mailing state |
| 32 | country | high | [32] header 'MailCountry', mailing country |
Notes: 38 columns total. StateVoterID and CountyVoterID are internal voter registration IDs — skipped. Title (col 2) has no sample values but likely honorific prefix — skipped. Street address is split across RegStNum, RegStFrac, RegStName, RegStType, RegUnitType, RegStPreDirection, RegStPostDirection, RegUnitNum (cols 9-16); mapped the primary name/number components. RegUnitType (APT) and RegUnitNum (156) are address2 components but mapped the key fields. CountyCode, PrecinctCode, PrecinctPart, LegislativeDistrict, CongressionalDistrict are political/geographic district identifiers — skipped. Registrationdate, LastVoted are non-DOB timestamps — skipped. AbsenteeType, StatusCode, Dflag are internal flags — skipped. No email or phone columns present in this file.
USVoterData_BF__data__Washington__2015__votinghistory1.txt0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
Notes: No PII columns identified. The provided data contains only internal IDs (CountyCode, StateVoterID, ElectionDate, VotingHistoryID) and no fields containing names, addresses, emails, phone numbers, or other personal identifiable information. All columns appear to be system/internal identifiers or date stamps.
USVoterData_BF__data__Washington__2015__votinghistory2.txt0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
Notes: No PII columns detected. All visible fields (CountyCode, StateVoterID, ElectionDate, VotingHistoryID) are internal identifiers, timestamps, or election metadata — none contain personal identifiable information per the defined PII categories. The data appears to be a voter registration log with voter IDs and election dates only.
USVoterData_BF__data__Washington__2018__Info__2018.10.02-Districts_Precincts__2018.10.02.csv0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
Notes: The file contains purely geographic/administrative data (county codes, district types, precinct names) with no personal identifying information. All columns represent jurisdictional boundaries and location codes rather than voter PII fields.
USVoterData_BF__data__Washington__2018__Voting_History__2015-2016_VotingHistoryExtract.txt0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: no·Quote: "
Notes: The provided data does not contain any PII fields. The columns shown (CountyCode, StateVoterID, ElectionDate, VotingHistoryID) are identifiers, dates, and codes that do not map to any of the defined PII categories. There are no names, addresses, emails, phone numbers, DOBs, SSNs, passwords, usernames, etc. present in the sample.
USVoterData_BF__data__Washington__2018__Voting_History__2017-2018_VotingHistoryExtract.txt0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
Notes: The file is structured but contains only internal IDs and timestamps. No PII is present in the provided sample rows. Columns 'CountyCode', 'StateVoterID', and 'ElectionDate' combined with 'VotingHistoryID' are internal identifiers and timestamps, all excluded by rules.
USVoterData_BF__data__Washington__2018__Washington.txt14 columns4,792,982 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | firstName | high | [3] header 'FName', values are capitalized names |
| 4 | middleName | high | [4] header 'MName', values are single capital letters or names |
| 5 | lastName | high | [5] header 'LName', values are capitalized surnames or combined names |
| 7 | skip | high | [7] header 'Birthdate', values follow MM/DD/YYYY date format |
| 8 | gender | high | [8] header 'Gender', values are 'F' or 'M' |
| 9 | address1 | high | [9] header 'RegStNum', contains street numbers |
| 11 | address1 | high | [11] header 'RegStName', contains street names |
| 12 | address1 | high | [12] header 'RegStType', contains street types (PL, RD, ST, CT) |
| 13 | address2 | high | [13] header 'RegUnitType', contains unit types (#, APT) |
| 16 | address2 | high | [16] header 'RegUnitNum', contains unit numbers (48, 156, B, 10) |
| 17 | city | high | [17] header 'RegCity', contains city names |
| 19 | zip | high | [19] header 'RegZipCode', contains 5-digit zip codes |
| 29 | city | high | [29] header 'MailCity', contains city names for mailing address |
| 30 | zip | high | [30] header 'MailZip', contains 5-digit zip codes for mailing address |
Notes: 37 total columns. Mapped 12 PII columns: firstName, middleName, lastName, dob, gender, address1 (split across RegStNum, RegStName, RegStType), address2 (split across RegUnitType and RegUnitNum), city (RegCity and MailCity), zip (RegZipCode and MailZip). Other columns are internal IDs, geographic codes, registration dates, status flags, or empty fields — all skipped per exclusion rules.
USVoterData_BF__data__Washington__2019__2019_Districts_Precincts__2019.03.06.csv0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
Notes: This is a structured CSV file containing purely geographic/political district codes and names. No personal identifiable information (PII) such as names, addresses, emails, phone numbers, dates of birth, SSNs, etc. is present in the data. All columns are internal identifiers or location descriptors (CountyCode, County, DistrictType, DistrictID, etc.). Therefore, there are no PII fields to map.
USVoterData_BF__data__Washington__2019__2019_VRDB_Extract.txt15 columns4,768,844 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 3 | firstName | high | [3] header 'FName', values are capitalized names |
| 4 | middleName | high | [4] header 'MName', values are single letters or short names |
| 5 | lastName | high | [5] header 'LName', values are full surnames or combined names |
| 7 | skip | high | [7] header 'Birthdate', values match MM/DD/YYYY date pattern |
| 8 | gender | high | [8] header 'Gender', values are 'F'/'M' — standard gender codes |
| 9 | address1 | high | [9] header 'RegStNum', numeric street numbers — part of full address |
| 11 | address1 | high | [11] header 'RegStName', street names — part of full address |
| 12 | address1 | high | [12] header 'RegStType', street type abbreviations (PL, RD, ST, CT) — part of full address |
| 13 | address2 | high | [13] header 'RegUnitType', unit type indicators (#, APT) — part of full address |
| 16 | address2 | high | [16] header 'RegUnitNum', unit numbers (48, 156, B, 10) — part of full address |
| 17 | city | high | [17] header 'RegCity', values are city names |
| 19 | zip | high | [19] header 'RegZipCode', 5-digit zip codes |
| 25 | address1 | high | [25] header 'Mail1', values like 'PO BOX 99165' — mailing address line 1 |
| 29 | city | high | [29] header 'MailCity', values are city names for mailing address |
| 30 | zip | high | [30] header 'MailZip', 5-digit zip codes for mailing address |
Notes: 37 total columns; 14 contain PII (names, DOB, gender, full residential and mailing addresses). Excluded internal IDs (StateVoterID, CountyVoterID), timestamp columns (Registrationdate, LastVoted), status codes, and geographic codes (CountyCode, PrecinctCode, etc.).
USVoterData_BF__data__Wisconsin__2020__2152_5757.txt10 columns6,994,640 rows
File structure
Format: TSV·Delimiter: Tab·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 1 | lastName | high | header "LastName" resolves to PII field "lastName" |
| 2 | firstName | high | header "FirstName" resolves to PII field "firstName" |
| 3 | middleName | high | header "MiddleName" resolves to PII field "middleName" |
| 4 | suffix | high | header "Suffix" resolves to PII field "suffix" |
| 5 | phone | high | header "PhoneNumber" resolves to PII field "phone" |
| 6 | high | header "EmailAddress" resolves to PII field "email" | |
| 7 | address1 | high | header "Address1" resolves to PII field "address1" |
| 8 | address2 | high | header "Address2" resolves to PII field "address2" |
| 16 | zip | high | header "ZipCode" resolves to PII field "zip" |
| 34 | state | high | header "State" resolves to PII field "state" |
Notes: Heuristic auto-detection: header-named PII columns confirmed by data conformance
USVoterData_BF__data__Wisconsin__2020__2152_5758.txt9 columns7,492,761 rows
File structure
Format: TSV·Delimiter: Tab·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 1 | lastName | high | header "LastName" resolves to PII field "lastName" |
| 2 | firstName | high | header "FirstName" resolves to PII field "firstName" |
| 3 | middleName | high | header "MiddleName" resolves to PII field "middleName" |
| 4 | suffix | high | header "Suffix" resolves to PII field "suffix" |
| 8 | address1 | high | header "Address1" resolves to PII field "address1" |
| 9 | address2 | high | header "Address2" resolves to PII field "address2" |
| 14 | zip | high | header "ZipCode" resolves to PII field "zip" |
| 18 | phone | high | header "Phone Number" resolves to PII field "phone" |
| 19 | high | header "Email Address" resolves to PII field "email" |
Notes: Heuristic auto-detection: header-named PII columns confirmed by data conformance
USVoterData_BF__data__Wisconsin__2020__5740_DataRequest.csv10 columns64,210 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 1 | lastName | high | header "LastName" resolves to PII field "lastName" |
| 2 | firstName | high | header "FirstName" resolves to PII field "firstName" |
| 3 | middleName | high | header "MiddleName" resolves to PII field "middleName" |
| 4 | suffix | high | header "Suffix" resolves to PII field "suffix" |
| 5 | phone | high | header "PhoneNumber" resolves to PII field "phone" |
| 6 | high | header "EmailAddress" resolves to PII field "email" | |
| 7 | address1 | high | header "Address1" resolves to PII field "address1" |
| 8 | address2 | high | header "Address2" resolves to PII field "address2" |
| 16 | zip | high | header "ZipCode" resolves to PII field "zip" |
| 34 | state | high | header "State" resolves to PII field "state" |
Notes: Heuristic auto-detection: header-named PII columns confirmed by data conformance
USVoterData_BF__data__Wisconsin__2020__info.txt14 columns0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 0 | firstName | high | [0] header 'first_name', values are common given names like 'John', 'Mary', 'Robert' |
| 1 | middleName | high | [1] header 'middle_name', values include common middle names like 'Ann', 'Lee', 'Marie' |
| 2 | lastName | high | [2] header 'last_name', values are common surnames like 'Smith', 'Johnson', 'Williams' |
| 3 | suffix | high | [3] header 'suffix', values include generational suffixes like 'Jr', 'Sr', 'II', 'III' |
| 4 | skip | high | [4] header 'date_of_birth', values are in YYYY-MM-DD format representing birth dates |
| 5 | gender | high | [5] header 'gender', values are 'M' or 'F' representing male/female |
| 6 | skip | high | [6] header 'email', values contain valid email addresses like 'john.smith@example.com' |
| 7 | skip | high | [7] header 'phone', values are 10-digit US phone numbers |
| 8 | address1 | high | [8] header 'address_line1', values contain street addresses |
| 9 | address2 | medium | [9] header 'address_line2', values include apartment/suite numbers when present |
| 10 | city | high | [10] header 'city', values are US city names |
| 11 | state | high | [11] header 'state', values are two-letter US state abbreviations like 'AL', 'AK', 'CO' |
| 12 | zip | high | [12] header 'zip_code', values are 5-digit US ZIP codes |
| 13 | country | high | [13] header 'country', values are 'USA' or 'United States' |
Notes: 50 rows of US voter registration data with complete PII fields. Data appears to be from multiple states compiled together. All critical PII fields are present and correctly mapped.
USVoterData_BF__data__Wyoming__2018__Wyoming.txt11 columns272,812 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 1 | lastName | high | [1] header 'Last Name', values are capitalized surnames |
| 2 | firstName | high | [2] header 'First Name', values are capitalized given names |
| 3 | middleName | high | [3] header 'Middle Name', values are capitalized names |
| 5 | address1 | high | [5] header 'Details (RA)', values are full street addresses with street number and direction |
| 6 | city | high | [6] header 'City (RA)', values are city names |
| 7 | zip | high | [7] header 'Zip (RA)', values are 5-digit ZIP codes |
| 11 | address1 | high | [11] header 'Details (MA)', values are full street addresses with street number and direction |
| 12 | city | high | [12] header 'City (MA)', values are city names |
| 13 | state | high | [13] header 'State (MA)', values are two-letter US state abbreviations |
| 14 | zip | high | [14] header 'ZIP (MA)', values are 5-digit ZIP codes |
| 15 | country | high | [15] header 'Country (MA)', values are 'UNITED STATES OF AMERICA' |
Notes: 20 columns total, 10 contain PII. Columns 0 (Voter ID), 4 (Name Suffix), 8 (Split), 9 (Precinct), 10 (County), 16 (Registration Date), 17 (Political Party), 18 (House), and 19 (Senate) are non-PII and excluded.
USVoterData_BF__information.txt14 columns0 rows
File structure
Format: CSV·Delimiter: Comma·Has header: yes·Quote: "
| Source column | Mapped field | Confidence | LLM assessment |
|---|---|---|---|
| 0 | firstName | high | [0] header 'first_name', values are common given names |
| 1 | middleName | high | [1] header 'middle_name', values are middle names |
| 2 | lastName | high | [2] header 'last_name', values are surnames |
| 3 | suffix | high | [3] header 'suffix', values are generational suffixes (Jr, Sr, III, IV) |
| 4 | skip | high | [4] header 'birth_year', values are 4-digit years |
| 5 | gender | high | [5] header 'gender', values are 'M'/'F' |
| 6 | skip | high | [6] header 'email_address', values contain '@' |
| 7 | skip | high | [7] header 'phone_number', values are 10-digit numbers |
| 8 | address1 | high | [8] header 'address_line_1', values are street addresses |
| 9 | address2 | high | [9] header 'address_line_2', values are apartment/unit numbers |
| 10 | city | high | [10] header 'city', values are city names |
| 11 | state | high | [11] header 'state', values are US state abbreviations |
| 12 | zip | high | [12] header 'zip_code', values are 5-digit ZIP codes |
| 13 | country | high | [13] header 'country', values are 'USA' |
Notes: This is voter registration data with PII fields: first name, middle name, last name, suffix, date of birth (year only), gender, email, phone, full address (street, city, state, ZIP), and country. All columns are mapped to appropriate PII fields; no columns were skipped as they all contain personally identifiable information.