5 PASS SUF Datasets and Variables: Collection, Editing and Usage
Contact authors: Beste, Jonas; Chauhan, Madhu; Frodermann, Corinna; Dummert, Sandra; Malich, Sonja; Prospero, Valentina; Wenzig, Claudia.
The data collected in each wave of PASS is anonymised and made available to the scientific community for secondary analysis in the form of a Scientific Use File (SUF). The raw survey data collected in each wave are processed through a series of editing, coding, and documentation steps to ensure consistency, accuracy, and comparability over time.
The SUF contains a rich combination of household- and individual-level information, covering both cross-sectional and longitudinal components, along with episode-based (“spell”) data.
In addition to the survey responses, the SUF incorporates structural identifiers, weighting factors, and metadata required to link datasets across waves. Its design allows researchers to conduct both descriptive and analytical studies, ranging from short-term behavioural assessments to long-term panel analyses of households and individuals in Germany.
This chapter provides a detailed guide to understanding and working with the SUF. There are four major sections within this chapter. Section 5.1 describes the structure of the SUF, including core and temporary9 datasets, their levels, types, and formats, as well as the role of key and pointer variables in linking records. Section 5.2 explains the organisation of variables, covering naming conventions, classification into system and surveyed variables, and the use of standardised codes. Section 5.3 outlines the data-editing process carried out prior to the SUF release, including household and individual structure checks, filter verification, plausibility testing, and the integration of episode-based information. Finally Section 5.4 introduces the linkage of PASS data with administrative records and additional datasets such as PASS-ADIAB.
Together, these sections aim to equip users with the necessary knowledge to navigate the SUF efficiently, merge datasets accurately, and apply them in robust cross-sectional and longitudinal research.
5.1 Datasets in the SUF
5.1.1 Classification of Datasets in the SUF
The Scientific Use File (SUF) of PASS comprises multiple datasets that differ in content, frequency of collection, level, type, and format. A subset of these datasets is produced regularly in each survey wave; these are referred to as core datasets in this section. Core datasets form the backbone of the SUF and contain information collected consistently across waves, enabling robust longitudinal analyses.
In contrast, temporary datasets refer to those collected only for specific periods or for research needs and are no longer part of the current SUF release. However, they were briefly included within the previous SUF releases and hence briefly discussed.
An updated graphical overview of the dataset structure can be found in wave-specific data reports; for example, see Figure 2: Dataset Structure of the PASS in the PASS Datenreport Wave 17 (Anker et al., 2025).
Table 5.1 provides an overview of the core datasets and a detailed classification of these datasets based on their level, type, and format, as well as their key and pointer variables.
This table classifies the core PASS datasets by filename, level, type, format, key variables, and pointer variables.| Dataset (filenames) | Level | Type | Format | Key variables | Pointer variables |
|---|---|---|---|---|---|
| Household dataset (“HHENDDAT.dta”) | Household level | Cross section | Long |
(1) hnr (Household number); (2) welle (Indicator for survey wave) |
(1) uhnr (Original household number) |
| Person dataset (“PENDDAT.dta”) | Individual level | Cross section | Long |
(1) pnr (Constant person ID number); (2) hnr (Household number); (3) welle (Indicator for survey wave) |
(1) uhnr (Original household number) |
| Household register (“hh_register.dta”) | Household level | Register | Wide |
(1) hnr (Household number); (2) hnr$ (Household number in wave $) |
(1) uhnr (Original household number); (2) pnrzp$ (Constant personal ID number of person who gave the household interview in wave $) |
| Person register (“p_register.dta”) | Individual level | Register | Wide |
(1) pnr (Constant personal ID number); (2) hnr$ (Household number in wave $); (3) zplfd$ (Serial number of the target person in the household in wave $) |
(1) uhnr (Original household number); (2) zmhh$ (Constant personal ID number of target person’s mother living in the same household in wave $); (3) zvhh$ (Constant personal ID number of target person’s father living in the same household in wave $); (4) zparthh$ (Constant personal ID number of target person’s partner living in the same household in wave $) |
| Household weights (“hweights.dta”) | Household level | Weights | Long |
(1) hnr (Household number); (2) welle (Indicator for survey wave) |
No pointer variable |
| Person weights (“pweights.dta”) | Individual level | Weights | Long |
(1) pnr (Constant person ID number); (2) welle (Indicator for survey wave) |
No pointer variable |
| Citizen’s benefit spells (“alg2_spells.dta”) | Household level | Spells | Spell |
(1) hnr (Household number); (2) spellnr (Spell number) |
No pointer variable |
| Biography spells (“bio_spells.dta”) | Individual level | Spells | Spell |
(1) pnr (Constant person ID number); (2) spellnr (Spell number) |
No pointer variable |
| Children dataset (“KINDER.dta”) | Individual level | Cross-section | Long |
(1) pnr (Constant person ID number); (2) hnr (Household number); (3) welle (Indicator for survey wave) |
(1) zmhh (Constant personal ID number of mother living in same household in respective wave); (2) zvhh (Constant personal ID number of father living in same household in respective wave) |
| One-Euro-Jobs spells (“ee_spells.dta”) | Individual level | Spells | Spell |
pnr (Constant person ID number); (2) spellnr (Spell number) |
No pointer variable |
Before describing the datasets in detail, the following section explains the columns in the Table 5.1:
The first column lists all the core dataset names along with their corresponding file names as observed in the SUF (in parentheses).
Level classification, in the second column, distinguishes whether a dataset contains information at the household level or the individual level. Household-level datasets capture information about the household, collected through a structured household questionnaire administered to a knowledgeable household member at the beginning of the interview process. Individual-level datasets include data on persons aged 15 and above living in these households, obtained through separate personal interviews. Since each respondent is linked to a specific household in every survey wave, the two levels are inherently connected.
Type classification, shown in the third column of the Table 5.1, refers to the nature and content of the datasets contained in the SUF. Across both household and individual levels, four main dataset types are distinguished: cross-sectional, register, weighting, and spell datasets.
Cross-sectional datasets contain survey data collected during each household or individual interview, representing the situation at a specific point in time. They exclude the parts where the respondent was asked to report episodes (e.g., receipt of Citizen’s Benefit10), which are instead stored in spell datasets.
Spell datasets store episode-based information in which respondents report activities or events in the form of episodes. The respondent specifies a period starting in the past and reports all relevant activities or events up to the date of the interview. This way of collecting data particularly differs from the cross-sectional concept described above; therefore, it cannot be integrated directly in the cross-sectional datasets. Each reported episode forms a separate observation, includes a start date and end date, contains further information about its content, and allows multiple entries per respondent. Episodes may or may not overlap with the fieldwork period.
Register datasets provide basic structural information on survey participation and wave-specific identifiers. The household register lists all households ever surveyed in PASS, while the person register contains all individuals within these households. These datasets include core identifiers and participation histories, serving as a reference framework for linking data across waves.
Because of the complex sample design of PASS, weightings are necessary for all descriptive analyses. Hence, the SUF includes household- and individual-level weighting datasets that correspond structurally to the cross-sectional datasets. These contain weights for each wave in which a household or individual was surveyed, allowing researchers to project the sample data onto the target populations.
Format classification, shown in the fourth column of the Table 5.1, indicates how the data are structured within each dataset. In the SUF, datasets are prepared in one of three formats: wide, long, or spell.
In the wide format, each unit (household or individual) is represented by exactly one observation (row) in the dataset. Wave-specific information is stored in separate variables (columns) ending with the corresponding wave number. For waves in which no information is available for a unit, these variables are filled with standardised missing- value codes. The register datasets of PASS follow this format, meaning that each observation uniquely represents a specific unit and can be identified using a single key variable.
In the long format, each observation represents a specific wave in which the unit was surveyed, resulting in as many rows per unit as the number of waves it participated in. Wave-specific information appears as additional rows for the same unit. Cross-sectional and weighting datasets use this format. Variables remain consistent across waves, with each variable appearing only once as a column. Changes in question wording or survey design may result in the introduction of new variables, while variables that are surveyed only in specific waves are assigned the missing code “-9” in waves where they were not collected. Observations in long-format datasets are identified using a combination of key variables for the unit and the wave.
The spell format is used exclusively for episode-based data. In this format, each reported episode corresponds to a separate observation. Episodes may span multiple waves, and updated information can be added in subsequent waves if the episode is ongoing. Units with no reported episodes are not represented in spell datasets, while those with multiple episodes have one row per episode. Observations in spell datasets represent specific episodes of specific units and are uniquely identified using a combination of key variables for the unit and the spell number.
All datasets include key variables, which are identifiers used to distinguish observations within and across datasets of the SUF. Therefore, the fifth column lists the key variables of the datasets. They allow users to reliably link the same unit - household or individual; across different datasets of the SUF and, where applicable, across survey waves. For example, the household number (hnr) identifies a specific household, while the personal identification number (pnr) uniquely identifies an individual. In long-format datasets, key variables are typically combined with the wave indicator (welle) to ensure that each observation is uniquely identified. In spell datasets, the spell number (spellnr) is included alongside the unit’s key variable to distinguish between multiple episodes for the same unit.
A second group close to the key variables is the pointer variables, which are displayed in the last column. While key variables identify the same unit and link it between datasets, pointer variables establish links between different, related units. For example, the original household number (uhnr) can be used to link a split-off household to its household of origin. Similarly, in the person register, pointer variables such as the personal ID numbers of a respondent’s mother (zmhh$), father (zvhh$), or partner (zparthh$) can be used to map family relationships within and across households over time. Compared with the other datasets, the spell datasets do not contain any pointer variables. These variables enable complex data linkages that are essential for analysing household composition changes, kinship structures, and social relationships from a longitudinal perspective.
5.1.2 Description of Datasets by levels
Household-level Datasets
The household dataset (HHENDDAT.dta) is a cross-sectional dataset in long format. It contains data collected at the household level, with each row representing a household surveyed in a specific wave. Key variables include the household number (hnr) and the survey-wave indicator (welle). The original household number (uhnr) serves as a pointer variable for tracking household changes over time.
The household register (hh_register.dta) is a register dataset in wide format. It provides an overview of all surveyed households across waves, with wave-specific household numbers (hnr$) stored in separate columns. Key variables include the household number (hnr) and its wave-specific variations. Pointer variables include the original household number (uhnr) and the personal ID number of the respondent who provided the household interview in each wave (pnrzp$).
The household weights dataset (hweights.dta) contains weighting information required for analysis due to the complex sample design of PASS. This dataset, in long format, provides household weights for each surveyed wave. Key variables include the household number (hnr) and the survey-wave indicator (welle). No pointer variables are included.
The citizen’s benefit spells dataset (alg2_spells.dta) is a spell dataset in spell format. It records episodes of household-level Citizen’s Benefit receipt. Each observation represents a separate spell, identified by the household number (hnr) and the spell number (spellnr). No pointer variables are included.
Individual-Level Datasets
The person dataset (PENDDAT.dta) is a cross-sectional dataset in long format. It contains individual-level survey data, with key variables including the constant personal ID number (pnr), the household number (hnr), and the survey wave indicator (welle). The original household number (uhnr) is included as a pointer variable.
The person register (p_register.dta) is a register dataset in wide format, tracking individuals across survey waves. Each row represents a person, with key variables such as the personal ID number (pnr), wave-specific household numbers (hnr$), and the serial number of the target person in the household (zplfd$). Pointer variables include the original household number (uhnr) and the personal ID numbers of household members related to the target person, such as the mother (zmhh$), father (zvhh$), and partner (zparthh$).
The person weights dataset (pweights.dta) provides weighting information for individuals in long format. It includes the personal ID number (pnr) and the survey-wave indicator (welle). No pointer variables are included.
The biography spells dataset (bio_spells.dta) is a spell dataset in spell format, capturing individual life-course episodes collected since wave 2. Each observation represents a reported spell, with key variables including the personal ID number (pnr) and the spell number (spellnr). No pointer variables are included.
The children dataset (KINDER.dta) is a cross-sectional dataset in long format, containing information on children within surveyed households since wave 6. Key variables include the personal ID number (pnr), the household number (hnr), and the survey-wave indicator (welle). Pointer variables include the personal ID numbers of the child’s mother (zmhh) and father (zvhh), if they reside in the same household during the respective wave.
The One-euro-job spells dataset (ee_spells.dta) is a spell dataset at the individual level. It records episodes in which a person participated in or received an offer for a One-euro-job. Key variables include the personal ID number (pnr) and the spell number (spellnr). No pointer variables are included.
5.1.3 Additional dataset for Websurvey mode
In addition to CAPI and CATI, the web survey in PASS was conducted for the first time in wave 16 as an additional interview between two regular panels waves and is available as a single dataset named “Web.dta”. One rationale for introducing the web survey was to test optimal procedures for a future use of the web mode as a regular data collection mode for PASS. In addition, it was used for additional questionnaire modules that did not get priority in the core questionnaire.
Regardless of the wave in which the web survey was implemented, this dataset is cross-sectional at the individual level. The key variables are pnr (Constant person ID number), hnr (Household number) and welle (Indicator for survey wave). The websurvey dataset has no pointer variables.
More information about the datasets of the websurvey can be found in the corresponding data reports which can be downloaded from the FDZ homepage11.
5.1.4 Temporary datasets
In addition to the datasets collected annually, PASS includes several temporary datasets that were collected only once or, at most, twice. These include vignette datasets and spell datasets at the individual level.
Vignette datasets are identified by “VIGDAT” in the dataset name. There are currently three such datasets:
Job acceptance readiness (VIGDAT_SUB.dta) – collected only in wave 5.
Assessment of job offers (VIGDAT_KON.dta) – collected only in wave 12.
Mothers in employment (VIGDAT_MUK.dta) – collected only in wave 15.
The key variables for vignette datasets are typically pnr (constant personal ID number) and vignr (vignette number). They are generally presented in long format, and there are no pointer variables.
At the individual level, there are three temporary spell datasets: Unemployment Benefit I spell dataset and two active labour market programme datasets. The alg1_spells.dta dataset, collected only in wave 1, contains episodes of Unemployment Benefit I receipt, including start and end dates and the total amount of benefits per month. The two active labour market programme datasets are massnahmespells.dta (wave 1) and mm_spells.dta (waves 2 and 3), both of which record episodes of participation in specific employment or training measures. All three temporary spell datasets share the same key variables: pnr (constant personal ID number) and spellnr (spell number). None contain pointer variables.
In addition to these spell and vignette datasets, there is an individual-level dataset on retirement provision (PAVDAT.dta). This dataset, in long format, provides detailed individual-level information on retirement provisions. It was collected only in wave 3 and only for persons aged 40 to 64, or those with a partner in this age range. The key variables are pnr (constant personal ID number) and welle (indicator for survey wave). No pointer variables are included.
At the household level, there is a corresponding retirement-provision dataset (HAVDAT.dta), also collected only in wave 3. It contains detailed household-level information on retirement provisions, collected under the same age-related criteria as PAVDAT.dta. This dataset is in long format, with hnr (household number) and welle as key variables.
Finally, two datasets were collected only at the individual level in wave 1: Proxydata and Refusing Individuals. Further details on these temporary datasets can be found in data report for wave 1.
More information about the datasets of the temporary datasets is provided in the corresponding data reports or in the older versions of PASS User Guide (Bethmann et al., 2013, pp. 29–47), (see also Section 5.2).
5.2 Variables in PASS
Two principal approaches exist for naming variables in survey datasets. The first approach assigns variable names according to their respective order in the questionnaire. While this facilitates the identification of variables with their corresponding items within the questionnaire, it complicates longitudinal analyses and tracking across waves, as the sequence of questions often changes between waves.
The second approach, which has been adopted in PASS, assigns consistent variable names across waves, irrespective of their position in the questionnaire, supplemented by a wave indicator when necessary. This ensures that identical items retain the same variable name over time, thereby greatly facilitating longitudinal tracking and analysis. Although this reduces the questionnaire’s usefulness as a direct documentation tool12, the advantages, such as improved data management, comparability, and analytic efficiency in a long-term panel study, decisively outweigh this limitation. Moreover, the organisation of PASS data in long format makes the use of uniform variable names indispensable.
The codebook distinguishes between three different types of variables:
5.2.1 System variables
System variables are variables created during the survey process. They can be used, firstly, to comprehend the filters documented in the questionnaire. At least some of the system variables can also be of interest from a content-related or methodological point of view, for example the interview mode or the number of children of a certain age group living in the household. System variables are allocated individual names, for which lower-case letters and numbers are combined in some cases. The system variables also include the weights.
5.2.2 Surveyed variables
Surveyed variables are variables that were collected in this form directly in the questionnaire. These variables are given entirely new, abstract variable names. The concept behind this naming process is illustrated in the figure below using an example.
Figure 5.1: Variable naming scheme
\(A.\) The first letter of the variable name indicates the questionnaire level, i. e. household or individual dataset, and is represented accordingly by the uppercase letter H or P.
\(B.\) This is followed by one or two upper-case letters which indicate the subject area to which the variable belongs (see Table 5.2 for a complete list).
The datasets which are processed in spell form, contain no introductory P or H variables. Instead, the variables in these datasets are given a uniform subject-based name consisting of two or three letters or two letters and one digit (e.g. AL20100, EE0100a; see Table 5.3).
\(C.\) The introductory letter combination is then followed by two consecutively allocated numbers, which indicate the number of the question within the subject area.
\(D.\) The last two digits serve as placeholders for further specifications and are set to two zeros by default. However, they can be adjusted in specific cases. New variables introduced in later waves within a subject area can be assigned intermediate values (e.g. PAS0850 between PAS0800 and PAS0900). Additionally, this option has been used in cases where a second variant including coded information from an open-ended survey question or response category has been made available in addition to the original version of the variable (see Section 5.2.3). Furthermore, the placeholder is used to indicate wave-specific variables in spell data such as e. g. AL20804 (wave 4), AL20805 (wave 5), AL20806 (wave 6) and so on.
\(E.\) In the case of variables for items from multi-item batteries or in a looped sequence of questions, a further lower-case letter may be added to identify the item or the current cycle within the loop.
To clearly distinguish variables from the web survey, all variable names originating from the web survey are additionally prefixed with a lowercase “w.” Some of the items collected in the web survey are also included in the PENDDAT dataset from the main panel survey. However, the variables may differ in their response categories between the web survey and the main panel survey. For instance, the list of response categories for the question about practised sports (wPSB0100/wPSB0101 vs. PSB0100/PSB0101) has been both shortened and modified in the web survey (as compared to the main panel survey).
This table lists indicator codes used in cross-sectional PASS variable names along with their subject area indicators for both individual-level and household-level variables separately.| Individual code | Individual subject area | Household code | Household subject area |
|---|---|---|---|
| PA | General | HA | General |
| PAA | Educational aspiration | HBT | Education and Inclusion Subsidies |
| PAC | Labor market opportunities | HCV | Corona |
| PAS | Job-search | HD | Demography |
| PB | Education | HEK | Income |
| PCV | Corona | HKI | Child-care |
| PD | Demography | HLS | Standard of living |
| PEE | 1-Euro-Jobs | HR | Household economics and everyday economic practices |
| PEF | Dealing with finances | HT | Social participation |
| PEK | income | HW | Housing |
| PEO | Attitudes and orientations | ||
| PER | Employment & retirement | ||
| PET | Employment | ||
| PG | Health | ||
| PGR / PPG | Perception about justice | ||
| PKO | Cooperation plan | ||
| PLA | Latent functions of work | ||
| PLS | Standard of living | ||
| PME | Memory | ||
| PMI | Migration | ||
| PMJ | Mini Job | ||
| PML | Minimum wage | ||
| PMV | Vignette: Childcare and employment | ||
| PP | Care | ||
| PPT | Political participation | ||
| PQB | Quality of employment | ||
| PSB | Sports | ||
| PSH | Social origin | ||
| PSK | Social relations | ||
| PSM | Social media | ||
| wPSB | Web: Sports | ||
| PSU | Obligation for job search | ||
| PSV | Stigma awareness | ||
| PTK | Contact to social security institutions | ||
| PTS | Internet use | ||
| PV | Trust game | ||
| PWB | Further training | ||
| wPEA | Web: Employment | ||
| wPEO | Web: Attitudes and orientations | ||
| wPMK | Web: Media consumption | ||
| wPQB | Web: Quality of employment | ||
| wPSA | Web: Sanctions (Unemployment Benefit II) |
| Individual code | Individual subject area | Household code | Household subject area |
|---|---|---|---|
| AL (wave 1 only) | Receipt of Unemployment Benefit I | AL2 | Receipt of Unemployment Benefit II |
| AL (from wave 2) | Spells of registered unemployment and receipt of Unemployment Benefit I since January 2005 | ||
| ALM (wave 1 only) | Employment and training measures | ||
| EE (from wave 4) | One-Euro-Job | ||
| ET (from wave 2) | Employment with earnings of more than € 400 per month since January 2005 | ||
| MN (waves 2–3) | Employment and training measures | ||
| LU (waves 2–3) | Other activities since January 2005 |
5.2.3 Generated variables
The generated variables in the strict sense are aggregated from various other variables, e. g. from open-ended and categorical income measures, or they are even more complex constructs such as equivalised household income or classifications for education (such as ISCED or Casmin) or status (e. g. EGP, ESEC). Generated variables in this strict sense are allocated individual names that are as clear and memorable as possible, in lower-case letters (e.g. hhtyp or migration).
Another group of generated variables includes those in which information from open-ended survey questions or response categories were added to another (closed) variable. Although these variables are, strictly speaking, also generated variables and are classified as such in the frequency tables of the codebook, they are not given clear names. Instead, their names are based on those of the original variable, while the final “0” is replaced by “1” (e. g. the original variable PA0100a and its coded counterpart PA0101a).
5.3 Data Editing in PASS
5.3.1 Overview of Data editing
The SUF of PASS is an outcome of a comprehensive and systematic data editing process. In each wave, the field institute collects raw data, which is then verified, coded for open-ended responses, and transformed into variables before being integrated into the SUF datasets. While the process is refined and adapted from wave to wave, the overall logic and sequence of steps have remained consistent over time. Wave-specific procedures are documented in the corresponding data reports (see, for example, Berg, Cramer, Dickmann, Gilberg, et al. (2024), for wave 17), but this section provides a general overview of the key steps and their sequence.
Data editing for the first two waves was conducted at the Institute for Employment Research (IAB). From wave 3 onwards, the Institut für Angewandte Sozialwissenschaft (infas), the new field institute for PASS, assumed responsibility for this task13. To avoid any unexpected changes in the procedure or inconsistencies in the SUF, several safeguards accompanied this transition.
First, the new contract with infas explicitly required adherence to the sequence and methodology of the data editing process as that of previous waves. Accordingly, infas received the relevant syntax files and datasets from wave 2, along with documentation for each processing step.
Second, infas carried out the editing process in continuous coordination with the IAB. Major decisions, such as handling problematic household structures or integrating spell datasets (such as the bio_spells dataset introduced in wave 4), were made in consultation with the IAB. In addition, IAB remained available for discussions and inquiries throughout the editing period.
Third, once the SUF for wave 3 was completed, IAB reviewed the final dataset, focusing on both its structure and content.
Besides the data editing, infas also performs initial transformations of the raw ASCII data from interviews to produce a range of internal datasets, including:
a household dataset for cross-sectional questions (e.g., childcare),
a household dataset for longitudinal modules (e.g. ALG II),
a dataset tracking household composition over time (matrix),
a dataset representing household relationships (relationship matrix),
a person/senior dataset for cross-sectional questions,
Two separate person-level longitudinal datasets (e.g. employment history and labour market policy measures),
a dataset containing open-ended responses across all survey components.
These datasets are then subjected to the formal and content-based checks described below before being incorporated into the SUF.
Apart from these institutional changes, the logic and sequence of the data editing process have remained consistent across waves. The process can be divided into the following steps:
This table lists the main procedural steps involved in editing PASS data, from household structure checks to final SUF dataset checks.| No. | Step of the procedure |
|---|---|
| 1 | Check of the household structure of re-interviewed households |
| 2 | Removal of problematic/incomplete interviews (household and/or individual level) |
| 3 | Integration of individual dataset and senior citizen’s dataset |
| 4 | Correction of the household structure of re-interviewed households |
| 5 | Filter checks at the household level |
| 6 | Construction of a household grid dataset and plausibility checks |
| 7 | Generation of the synthetic benefit units (see description of variables in wave-specific data reports) |
| 8 | Generation of new control variables based on household data after filter checks and household grid plausibility checks |
| 9 | Filter checks at the individual level |
| 10 | Coding of information from open-ended survey questions |
| 11 | Plausibility checks of the household and individual-level data (excluding spell data) |
| 12 | Preparation, plausibility checks and construction of the spell datasets |
| 13 | Simple variable generations |
| 14 | Complex variable generations |
| 15 | Generation of the data structure for the scientific use file (household dataset, individual dataset, register dataset) |
| 16 | Anonymisation |
| 17 | Final check of the SUF datasets |
5.3.2 Structure Checks
The first step involves verifying the household structure of re-interviewed households by comparing it with the structure reported in the previous wave. This step ensures the identification and, where necessary, correction of implausible or problematic changes in household composition, as well as mis-allocations of individual interviews within households. For consistent longitudinal analysis, individuals must retain their positions within the household across waves and remain uniquely identifiable. The same personal identification number must never be assigned to different individuals in different waves.
If the correct household composition cannot be determined, all interviews from the respective household in that wave are excluded from the SUF. If only a single individual interview has a mismatch without further inconsistencies in household structure, then only that interview is removed.
To detect such problematic cases, multiple internal consistency checks are conducted, assessing, for example, the plausibility of household changes and the accuracy of individual assignments. These cases are then reviewed through a formalised process between infas and IAB, with IAB making the final decisions. However, not all checks result in exclusions, as many potential inconsistencies are pre-emptively avoided through built-in validations in the survey software (e.g., preventing all target persons from leaving a household simultaneously, or requiring at least one remaining person aged 15 or older).
The wave-specific data reports (e.g., Berg, Cramer, Dickmann, Gilberg, et al. (2024)) provide further details on these structural checks. Additionally, the household register (variables hnettok*, hnettod*) and person register (pnettok*, pnettod*) indicate removed cases across waves. It should be noted that not all deleted interviews are traceable in the SUF due to the way in which the register files are constructed14.
In addition, incomplete interviews at both the household and individual levels are excluded from the SUF15. Households that do not meet PASS’s definition of a successfully surveyed household (see Table 5.5)16 are also excluded. These cases are not recorded in the register datasets, as they were never deemed eligible to begin with, unlike the removed interviews described above.
This table summarises household types in PASS and the interview requirements at the household and individual levels.| Type of household | Household level interview | Individual level interview(s) |
|---|---|---|
| New household (interviewed for the first time and drawn for the initial sample or a refreshment sample) | Yes (completed) | Yes (at least one completed) |
| Re-interviewed household (household already interviewed in a previous wave of PASS) | Yes (completed) | None required |
| New split-off household (interviewed for the first time and split-off from another household in PASS) | Yes (completed) | None required |
5.3.3 Filter Checks and Assignment of Standardised Codes
Every variable included in the SUF undergoes a thorough filter check. During this process, the system flags filter violations and assigns standardised missing codes. Following Table 5.6 presents an overview of the standardised codes used in PASS:
This table lists standardised missing codes used in PASS data and explains their meaning.| Code | Explanation |
|---|---|
| -1 | Don’t know. |
| -2 | Details refused. |
| -3 | Not applicable (filter) (question not asked due to filter). |
| -4 | Question mistakenly not asked (question should have been asked). |
| -5 | Question-specific code No. 1, only allocated as required. |
| -6 | Question-specific code No. 2, only allocated as required. |
| -7 | Question-specific code No. 3, only allocated as required. |
| -8 | Implausible value. |
| -9 | Item not administered in wave. |
| -10 | Item not administered in questionnaire version. |
These codes fall into the following categories:
Missing values due to respondent answers (“-1” = don’t know, “-2” = refused),
Filter-based missing values (“-3” = not applicable, “-4” = mistakenly not asked),
Question-specific codes (“-5” to “-7”),
Implausible answers (“-8”),
Items not part of the questionnaire/wave (“-9”, “-10”).
Except for implausible responses, which were flagged in a later step, all other missing-value categories are addressed during this phase. Variables in the raw datasets are checked sequentially, reflecting the order in which they were collected. In this process, the codes “-3” and “-4” are assigned accordingly.
A variable receives the code “-3” if it should not have been asked based on filter logic. Likewise, questions that were mistakenly asked despite filter conditions are corrected to “-3”17.
While erroneously collected data can be recoded as “-3”, it was not possible to retroactively supply missing answers. If a question was skipped despite being required by the filter, the missing code “-4” (“mistakenly not asked”) is applied.
It is important to note that “-4” codes may also result from retrospective corrections in household structure. For example, if a person was wrongly recorded as having moved out but was later reassigned to the household, missing household data for this person is marked with “-4”.
If a “-4” appears on a filter-relevant variable, then all subsequent questions dependent on it are also coded “-4”, provided they were not asked. If a follow-up was still triggered by an alternative filter path, its valid value remains.
Codes “-5” to “-7” are question-specific and include both special missing codes and valid categories (e.g., top-coded income values). Codes “-9” and “-10” are used for items not included in specific questionnaires or waves: “-9” denotes absence in a given wave, while “-10” covers differences between questionnaire versions (e.g., regular vs. senior or variant household questionnaires used in waves 1 to 3).
5.4 Record Linkage to Administrative Data
A distinctive feature of PASS is the possibility of linking survey data to administrative records, provided respondents give their consent for record linkage in each wave. The individual survey responses are linked to administrative data of the Institute for Employment Research (IAB) containing additional information on episodes of Unemployment Benefit I receipt, Citizen’s Benefit receipt, employment, job search, and participation in active labour-market programmes.
The Research Data Centre (FDZ) of the German Federal Agency at IAB produces the PASS-ADIAB dataset, which combines all PASS respondents with the relevant administrative data. This combined dataset enables both substantive research where administrative data serve as a supplementary source of information and methodological research, while administrative records are used as a validation source (Kreuter et al., 2010). Linking PASS-ADIAB to the PASS survey datasets is possible via the variable pnr (constant personal ID number).
For the waves 1 to 16, approximately 85.3% of people who responded to the person-level questionnaire and consented to linkage were successfully linked (Dummert & Sauer, 2024).
For further details on PASS-ADIAB, see Antoni & Bethmann (2019), which provides an overview of research opportunities, data sources, record linkage, and data structure, as well as the current methods and data reports available on the FDZ website (https://fdz.iab.de/unsere-datenprodukte/personen-und-haushaltsdaten/pass/). For technical details regarding the linkage methodology, refer to Bachteler (2008). Analyses of determinants of consent to record linkage and potential selection biases are available in Beste (2011).
As PASS-ADIAB contains sensitive data, specific user conditions apply. This integrated file is available to researchers for onsite use at the FDZ in Nuremberg or at one of the other FDZ locations in Germany, France, Spain, Italy, Poland, Luxembourg, the UK, the USA and Canada (see https://fdz.iab.de/en/data-access/on-site-use/ for details and locations) and via remote execution.
These datasets have been referred to as “discontinued” datasets in the older PASS data reports, method reports and other PASS related documents.↩︎
Unemployment Benefits II is used interchangeably with various other terminologies across this user guide and more widely across German welfare benefits research. For a detailed description refer 0.3.↩︎
Data reports are a bundle including the data report for the latest wave as well as all the other waves which can be downloaded from http://doku.iab.de/fdz/pass/FDZ-Datenreporte_PASS_EN.zip↩︎
To link questionnaire numbers with their corresponding fixed variable names, we provide a correspondence list. https://fdz.iab.de/en/pd_hd/panel-study-labour-market-and-social-security-pass-version-0623-v2/↩︎
The contract with the former field institute, TNS Infratest, was initially limited to three waves. As a result, fieldwork from wave 4 onwards was re-tendered. In the new call for proposals, the IAB decided to include the task of data editing starting with wave 3. Consequently, infas, as the new field institute from wave 4 on, also carried out the data editing for wave 3.↩︎
In PASS, the SUF’s register files are net files. The household register includes all households ever successfully surveyed. The person register contains all persons living in those households at the time of interview. Interviews removed from households or persons not included in these registers e.g., first-time interviews in refreshment samples are not visible in the SUF.↩︎
Thus, interviews that are terminated before completion do not appear in the SUF datasets.↩︎
As the criteria for a “successfully surveyed” household vary by household type, some households in the SUF appear without individual interviews in certain waves.↩︎
For instance, if detailed vocational training data was collected although the respondent indicated no vocational qualification, this information is replaced with “-3”.↩︎