What do we mean by an observational system? Why design an observational system? When should you design an observational system? Who should design an observational system? How do you design an observational system? A local community health center was starting a program to support regular physical activity among people with high blood pressure. The program had one main objective: to help participants engage in about 45 minutes of moderate aerobic activity at least four times a week over the course of six months. The hope was that regular physical activity would help participants manage their blood pressure, improve cardiovascular health, and support their overall well-being. A related goal was for participants to identify forms of physical activity they enjoyed and could continue after the formal program ended. The center recruited 50 people who wanted to participate. After appropriate health assessments, participants attended workshops about nutrition, high blood pressure, physical activity, injury prevention, and ways to incorporate movement into daily life. They also discussed different forms of activity—such as walking, bicycling, swimming, or other activities appropriate to their interests and abilities. Rather than requiring everyone to use the same exercise routine, the center encouraged participants to choose activities that worked for them while working toward the overall activity goal. Participants were asked to keep activity logs and to meet in small groups with health center staff once a month to review blood pressure, discuss progress and barriers, and receive support and guidance. The center then had to decide how to evaluate the program. One important question was whether participants were able to maintain their physical activity routines over the six-month period. Because participants exercised independently and at different times and places, staff could not directly observe every activity session. The center also wanted to understand whether participants experienced changes in blood pressure, well-being, physical activity habits, and other relevant health indicators. How could the center design a system to gather accurate, respectful, and useful information about what participants were doing and what changes were occurring? Once you have identified your evaluation questions and determined what information you need, you need a systematic way to collect that information. That is the purpose of an observational system. Like the center in the example, you may need several ways to examine behaviors, experiences, conditions, program activities, and changes over time. This section focuses on designing observational systems that provide useful information while respecting participants' privacy, dignity, and rights. What do we mean by designing an observational system? An observational system is a structured way of gathering information about a program—what is being implemented, how people are participating, what conditions are present, and what changes appear to be occurring. "Observation" may involve directly watching activities or conditions, but it can also include other systematic ways of documenting program processes and outcomes. The appropriate methods depend on the evaluation questions, the setting, the people involved, and ethical and privacy considerations. Direct observation. Direct observation involves systematically observing activities, behaviors, settings, or conditions firsthand. For example, if an initiative aims to increase use of a public park, observers might document how the park is used at different times of day, during different weather conditions, or during special events. Observation in public settings should still be planned carefully, with attention to privacy, research ethics, and whether consent or public notice is appropriate. In program settings, participants should generally understand when and why observation is taking place. Participant observation. A participant observer takes part in the setting or activity while also documenting what they experience and observe. In the park example, a neighborhood resident involved in the initiative might document patterns of park use while participating in everyday activities there. Participant observation can provide important contextual information because the observer experiences the setting from within it, but observers should be transparent about their role when appropriate and follow relevant ethical and privacy guidelines. Self-reports. Some experiences and behaviors cannot or should not be directly observed. People may be asked to describe their own experiences through interviews, journals, surveys, activity logs, questionnaires, or other forms of first-person reporting. Self-reports can provide information that is unavailable through external observation, including people's motivations, experiences, perceptions, and activities outside a program setting. Like all methods, self-reporting has limitations, so it may be useful to combine it with other appropriate sources of information when possible. Reports from others. An observational system may also include information from people who have relevant knowledge of the conditions or activities being evaluated. Depending on the program, these might include educators, health professionals, family members, community partners, service providers, or other people with direct and appropriate knowledge. Information should be collected in ways that protect confidentiality and respect participants' privacy and consent. Electronic or technology-assisted observation. Devices such as activity trackers, environmental sensors, cameras, audio recorders, or health-monitoring equipment can sometimes provide useful information. These tools should be used only when appropriate to the evaluation and with careful attention to informed consent, privacy, data security, accessibility, and how collected information will be stored and used. The fact that a technology can collect information does not necessarily mean that collecting it is appropriate. Tests and assessments. Depending on what is being evaluated, assessments may include academic measures, skills demonstrations, health screenings, validated questionnaires, or other tools. Assessments should be appropriate for the population and purpose, and their limitations should be considered when interpreting results. Public and administrative records. Census information, public health data, employment statistics, school or organizational records, and other administrative data may provide information about community-level conditions or outcomes. Evaluators should consider data quality, privacy requirements, who is represented or missing from the data, and whether the information is appropriate for the questions being asked. Products or results of behavior. Sometimes it is more practical or appropriate to examine the results of an activity rather than observing the activity itself. For example, an environmental initiative might measure changes in water quality, litter, or air pollution rather than attempting to observe every behavior that contributes to those conditions. A school-based health initiative might examine participation, student-reported well-being, access to healthy meals or physical activity, or other appropriate indicators rather than relying on body size as a primary measure of success. In addition to deciding which methods to use, an observational system should specify when, where, how often, by whom, and under what circumstances information will be collected. These decisions should be guided by the evaluation questions and by ethical, practical, cultural, and accessibility considerations. Consider whether you want to examine the process of the effort—how the program was planned and organized and whether implementation matched the intended approach. You may also want to examine what was actually delivered, who participated, how long activities lasted, what adaptations were made, and which parts participants found useful. Finally, consider which outcomes or changes you want to understand and how those outcomes can be measured responsibly. Designing an observational system requires thinking carefully about what you need to know and selecting methods that can provide that information accurately, consistently, ethically, and feasibly. We will discuss this process in more detail below. Why design an observational system? If you're serious about evaluation, there are several reasons to design a strong observational system: It can help you collect reliable information. A system that clearly defines methods, timing, procedures, and responsibilities makes information collected across observers, settings, and time periods more consistent and useful. It can help you collect the information you actually need. A well-designed system focuses data collection on the evaluation questions rather than gathering large amounts of information without a clear purpose. It can make consistent data collection more likely. When the people collecting information understand and agree on the procedures, observations are more likely to happen when and how they are intended. It can make analysis easier. Consistent observations can support both quantitative analysis, which examines numerical patterns, and qualitative analysis, which examines experiences, meaning, context, and interpretation. It can help you avoid fragmented evaluation. A structured system helps connect the information you collect directly to the questions you are trying to answer. It can strengthen confidence in your findings. Clearly documented and consistently applied methods make it easier to explain how conclusions were reached and to identify the limitations of the evidence. It can strengthen accountability with communities, funders, and policymakers. A thoughtful evaluation based on systematically collected information can help stakeholders understand what the program accomplished, what remains uncertain, and what should happen next. It can help you share promising practices responsibly. Strong evidence can help you explain which program components appear effective, for whom, and under what conditions, rather than assuming that an approach that worked in one setting will automatically work everywhere. It can provide useful information about what is working, what is not, and what should be adjusted. When should you design an observational system? An observational system includes the methods used to examine and document program processes, activities, experiences, and outcomes. Because it is central to evaluation, the ideal time to design the system is before implementation begins. Doing so allows you to establish baseline information and monitor the program from the start. In practice, many community organizations begin formal evaluation after a program is already underway because of limited time, staffing, funding, or evaluation capacity. Whenever evaluation begins, the observational system should be designed around the questions you are trying to answer and the information that can still be gathered accurately. It is worth taking the time to develop a system that is realistic, ethical, and useful. Whenever possible, observations should cover a full program cycle from beginning to end. Some programs do not operate in fixed cycles, in which case observation may focus on individual participants, particular activities, implementation periods, or changing community conditions. Records created before a formal evaluation begins—such as program notes, activity logs, meeting records, or participant feedback—may provide useful information. Before including them, determine whether they were collected consistently enough and for purposes compatible with the current evaluation. Information created for one purpose should not automatically be treated as equivalent to data collected under a defined observational protocol. Beginning observation late may also mean that important early changes or implementation decisions have already occurred. When feasible, begin collecting relevant information early enough to establish a baseline and understand how the program develops over time. Who should design an observational system? The Community Tool Box emphasizes participatory research and evaluation. An observational system is generally stronger when the people who will collect information, the people whose experiences are being represented, and those who will use the findings have meaningful opportunities to contribute to its design. In smaller community-based organizations, staff and volunteers may already have multiple responsibilities. The observational system should therefore be realistic enough to carry out with available time, staffing, technology, and other resources. Involving the people responsible for implementation can help identify unnecessary burdens and make the system more practical. The design process should include people who will actually collect the information, along with people who understand the program and the communities involved. Researchers or evaluators can contribute methodological expertise, while program participants and community members can help ensure that the system reflects community priorities, cultural context, accessibility needs, privacy concerns, and the realities of participation. The design team may therefore include members of the broader evaluation group as well as people specifically recruited to help develop the observational system. The group might include: Program staff, volunteers, and administrators Staff responsible for records, data management, technology, or administrative support External evaluators, researchers, or research consultants Program participants and community members affected by the work Community partners or trained volunteer observers If the design group does not include anyone with evaluation or research experience, participants may benefit from training or technical assistance on observational methods, ethics, privacy, measurement, and the kinds of information different methods can provide. If experienced researchers are involved, they can share this knowledge as part of a collaborative design process rather than assuming that methodological expertise should determine the system without community input. How do you design an observational system? Once you have developed your evaluation questions and overall evaluation plan, the next step is deciding how you will gather the information needed to answer those questions. Review your evaluation questions Return to what you originally wanted to understand about the program. In the health center example, the program was designed to support regular physical activity among people with high blood pressure, with participants working toward approximately 45 minutes of moderate activity four times a week over six months. The program hoped to contribute to several outcomes: Participants would experience improved blood pressure management Participants would increase or maintain regular physical activity in ways appropriate to their health, abilities, and preferences Participants would report improvements in overall well-being Participants would continue forms of physical activity they found sustainable after the six-month program ended Some evaluation questions might therefore include: To what extent were participants able to engage in their planned physical activity during the program? How did participants' blood pressure change over the six-month period? What changes did participants report in physical functioning, well-being, or confidence in maintaining regular activity? Other questions might include: How well attended were the workshops? Did participants find them useful, accessible, and relevant? Were particular program components associated with stronger engagement or outcomes? Did participants report changes in their overall well-being by the end of six months? Did participants continue engaging in regular physical activity after the formal program ended, and were improvements in blood pressure or well-being sustained? These may represent only some of the information the center wants. Your own evaluation may also examine planning, implementation, participant experience, accessibility, equity, timelines, adaptations, and intermediate benchmarks. Each of these questions should inform the design of the observational system. Decide what you need to observe to answer your questions The kinds of information you need will depend on the program, community context, and evaluation questions. Common areas include: Participants' behavior or activities. This might include participation in physical activity, demonstrated skills in an employment-training program, use of community resources, or interactions in a program setting. Behaviors should be defined carefully and observed in ways that respect participants' privacy and dignity. The behavior or practices of others involved. An evaluation might examine how staff interact with participants, how institutions implement a policy, or how other groups respond to an intervention. Examining organizational and staff practices can be particularly important when program outcomes depend on more than participant behavior alone. Conditions. An initiative may aim to change physical, social, environmental, organizational, or policy conditions—for example, improving affordable housing, restoring a polluted waterway, increasing accessibility, or changing institutional practices. Products or results of behavior or conditions. When an activity cannot or should not be observed directly, you may need to measure relevant outcomes or indicators. For example, a sexual health program might examine rates of sexually transmitted infections, access to preventive services, contraceptive use reported confidentially by participants, or other appropriate indicators rather than attempting to observe private behavior directly. When you rely on indirect indicators, make sure they are meaningfully connected to the behavior or condition you are interested in. Consider other factors that could influence the same outcome, and avoid treating one indicator as definitive evidence when multiple explanations are possible. Participants' knowledge, perceptions, or attitudes. These might be assessed through interviews, surveys, validated scales, knowledge assessments, or other appropriate methods. The knowledge, perceptions, or attitudes of other groups. An advocacy initiative, for example, might examine changes in public understanding, decision-maker awareness, organizational practices, or community support for a policy. Goal attainment. Some efforts focus on a specific outcome, such as adoption of a policy, creation of a community resource, or completion of an infrastructure project. Even when the final goal is clear, it can still be useful to document intermediate progress, relationships, barriers, and lessons learned. Interactions. An evaluation may focus on how people communicate, collaborate, share decisions, respond to one another, or experience relationships within a program. For example, a family-support effort might examine patterns of communication and engagement rather than simply counting interactions. These areas may relate either to program outcomes—what the program is intended to accomplish—or to process and implementation—how the work is planned and carried out. Areas related specifically to program process and implementation may include: Planning. Who participated in planning? How were decisions made? Whose perspectives were included or missing? How well did the process reflect community priorities? Timeline. When did planning, implementation, evaluation, and follow-up begin? Were timelines realistic? If delays occurred, what contributed to them? Participation and engagement. How many people participated? How long did they remain involved? Were there patterns in who participated, who did not, or who stopped participating? What barriers or supports affected engagement? Methods. What strategies, activities, or methods were used? Were they implemented as intended, and what adaptations were made? Program implementation. What activities actually occurred? How often and for how long? Who participated? Where did activities take place? What resources were used? What changed during implementation, and why? Whatever you choose to observe, define it carefully so that people collecting information understand what counts and what does not. Clear definitions, examples, and boundaries improve consistency and reduce the likelihood that different observers will interpret the same event in substantially different ways. Returning to the health center example, the center might use activity logs or confidential self-reports to understand how often participants engaged in physical activity. Blood pressure could be measured using appropriate clinical procedures. Participants could report changes in well-being, energy, confidence, or barriers to activity through surveys or interviews. Follow-up interviews or activity logs could help determine whether participants continued physical activity after the program ended. Decide how the observations will be conducted Earlier, we introduced several methods of observation. Each has strengths, limitations, and ethical considerations. Direct observation. Direct observation may involve an evaluator, staff member, trained community observer, or other designated person observing defined activities or conditions. In program settings, participants should generally know that observation is occurring and understand its purpose. Observation of public settings may sometimes be conducted without identifying individual people, but privacy, ethics, local expectations, and applicable requirements should still guide the design. Journals and activity logs can also support direct or participant observation. Observers may record what occurred, the context, relevant events, and their interpretations soon after an activity. When several people document the same period or activity, these records can provide multiple perspectives on how the program unfolded. The format of journals or logs should fit the program and the people using them. Written journals may work well in some settings, while voice notes, structured forms, digital logs, or other accessible formats may work better in others. Participant observation. Participant observers take part in the setting or activity while documenting what they experience. They may be community members, program participants, staff, or evaluators who participate openly in activities. Their position can provide valuable insight into relationships, context, and experiences that may not be visible to an external observer. For example, a small-grant program designed to support locally owned businesses might involve staff and community participants in workshops on budgeting, lending, business planning, and peer support. Staff and participants could reflect together on how the activities worked, which barriers emerged, and what participants found useful. In this setting, participant observation could contribute to learning while treating participants as partners in interpreting their own experiences rather than simply as subjects of observation. Self-reports. When relevant behaviors or experiences occur outside the program, participants may be the most appropriate people to describe them. Self-reports provide direct access to people's experiences, but responses can be influenced by memory, question wording, social expectations, concerns about privacy, or what participants believe the evaluator wants to hear. Clear questions, confidentiality protections, respectful data collection, and multiple sources of evidence can improve the usefulness of self-reported information. Participants may benefit from clear guidance about what information is useful to record, especially when completing activity logs or journals. At the same time, evaluators should avoid coaching people toward preferred responses. Self-reporting is especially valuable because it allows participants to describe experiences, motivations, barriers, and changes that an outside observer may not be able to see. Self-reports can be collected through individual and group interviews, focus groups, community meetings, surveys and questionnaires, journals, checklists, activity logs, or other ways for participants to share information. Reports from others. These may come from people who interact with participants or have relevant knowledge of the conditions being evaluated, such as educators, health professionals, service providers, family members, employers, or community partners. Consider whether the person has appropriate knowledge, whether participants have consented when necessary, and how the reporter's relationship, expectations, or biases may influence what they report. Electronic or technology-assisted observation. Technology can sometimes collect information that would be difficult to obtain otherwise. Examples include environmental sensors, activity monitors, clinical equipment, or traffic counters. Cameras, audio recording, location tracking, and other potentially intrusive technologies require especially careful consideration of informed consent, privacy, data minimization, data security, and who will have access to the information. Technology-generated data may appear objective, but equipment settings, algorithms, measurement error, and interpretation can still introduce limitations or bias. Tests or other assessment tools. Education, health, and human service organizations may use assessments to understand skills, knowledge, development, health status, or other outcomes. Whenever possible, use tools that are valid and appropriate for the population and purpose. Interpretation should also consider factors such as language, disability, culture, anxiety, fatigue, access to preparation, and testing conditions. Public and administrative records. Community-level indicators such as injury rates, health outcomes, school attendance, housing conditions, or transportation patterns may be available through government agencies, institutions, or other administrative systems. Consider the quality, timeliness, completeness, privacy protections, and limitations of these data before using them. In the health center example, blood pressure could be measured at regular intervals using appropriate clinical equipment and procedures. Participants' physical activity could be documented through self-reported activity logs, interviews, or optional activity-tracking tools if participants choose to use them. Well-being could be assessed through participant reports using appropriate questions or validated measures. Follow-up interviews, surveys, activity logs, or other appropriate methods could help determine whether participants continued regular physical activity after the program ended. Combining several types of information would provide a more complete picture than relying on any single measure. Decide when you need to observe Consider when information should first be collected and how often observations should occur. In many evaluations, collecting baseline information before or at the beginning of the program is important because it gives you a reference point for interpreting later changes. Possibilities include: Before-and-after observation. Information is collected at the beginning and end of an evaluation period. This can show whether change occurred, but by itself may provide limited information about when change happened, why it happened, or which aspects of the program contributed to it. For many evaluations, before-and-after measurement is most useful when combined with additional observations during implementation. Beginning with baseline information helps establish where participants, organizations, or community conditions started. Repeated measurement can then show whether change occurred gradually, rapidly, at particular stages, or only after the formal intervention ended. Even when an effort has one highly visible goal—such as adoption of a policy or completion of a community project—evaluating only whether the final goal was reached leaves out valuable information. Documenting strategies, relationships, challenges, intermediate accomplishments, and adaptations can help explain what contributed to the result and what can be learned for future efforts. At regular intervals during the evaluation period. Observations might occur hourly, daily, weekly, monthly, or at another regular interval, depending on what is being measured. Regular schedules make it easier to compare information across time. At selected or randomly chosen intervals. Some observations may occur at varying times to capture a broader range of conditions or reduce the likelihood that the same circumstances are observed repeatedly. Random or strategically varied sampling can sometimes provide a more representative picture. At specific times or stages. You may want to examine conditions at particular stages of implementation or under different circumstances. For example, park use might be observed on weekdays and weekends, during different seasons, at different times of day, or during community events. A program evaluation might gather information during planning, outreach, implementation, evaluation, and follow-up. Ongoing observation. In some settings, staff, participants, sensors, logs, or other systems may provide ongoing information. Continuous collection should be used only when it is useful and proportionate to the evaluation need. Particularly when recording people or collecting digital data, consider whether continuous monitoring is necessary and how privacy will be protected. At the health center, blood pressure might be measured at regular intervals during monthly meetings. Participants might document physical activity throughout the program using activity logs. Follow-up information could then be collected several months after the program to understand whether participants continued activities they found sustainable. Define and describe the behaviors, products, conditions, and/or events that observers should be concerned with Observers need clear definitions of what they are expected to document. The planning group—ideally including people who will conduct observations and people whose experiences are being represented—should establish shared definitions, examples, non-examples, and scoring or recording guidance for each element being observed. For example, an evaluation examining bullying or interpersonal aggression would need clear behavioral definitions rather than relying on assumptions about what those terms mean. Design training for observers Depending on their roles and previous experience, observers may need preparation in several areas: What should be recorded, and why. Observers should understand which contextual details matter, such as date, time, location, duration, environmental conditions, who was present, and other circumstances that might influence what occurred. Context can be as important as the event itself when interpreting observations. The definitions and descriptions of the behaviors, conditions, events, or situations being observed. Clear definitions improve consistency only if observers understand and practice using them. The effects of observation. People may behave differently when they know they are being observed. Observers should understand this possibility and use methods that reduce unnecessary influence while remaining transparent and respectful about the observation process. Audio or video equipment can also affect behavior. Participants should understand what is being recorded, why it is being recorded, how the recordings will be used, who will have access, and how long they will be retained. Appropriate consent should be obtained before recording begins. Observer bias. Observers' relationships, expectations, experiences, identities, professional roles, and assumptions can influence what they notice and how they interpret it. This is especially important when staff members are evaluating programs in which they are invested. Training, reflection, structured observation tools, multiple observers, and opportunities to discuss differences in interpretation can help reduce the influence of bias. The goal is not to assume that observers can eliminate all perspective, but to recognize and manage how perspective can shape observation. Observer drift. Over time, observers may gradually interpret definitions differently from how they were originally established. Periodic review, refresher training, and comparison across observers can help maintain consistency. Observational systems should include ways to identify and address observer effects, bias, inconsistent interpretation, or drift throughout the evaluation. Devise checks for reliability and accuracy Reliable information requires observers to apply definitions and measurement procedures consistently. Training is important, but it is also useful to periodically check whether different observers interpret the same behavior, condition, or event in similar ways. A participatory design process can strengthen this consistency because observers help define and understand what is being measured. Ways to strengthen agreement and consistency include: Use a clearly defined standard. Observation criteria should describe the behaviors, conditions, or indicators being assessed and provide enough detail for different observers to apply them consistently. Checklists, coding guides, rating criteria, examples, and non-examples can support this process. Research and evaluation teams often use standardized definitions and measurement procedures to improve consistency. Standards may include behavioral descriptions, clinical thresholds, environmental measurements, scoring rubrics, or other clearly defined indicators. Check inter-rater reliability. Inter-rater reliability refers to the degree to which different observers assess the same event or information consistently. Two or more observers may independently rate the same examples or observations and then compare results. The appropriate level of agreement depends on the type of measure and evaluation design; a single percentage threshold should not automatically be treated as adequate for every situation. If observers disagree, examine whether definitions, training, scoring procedures, or assumptions need clarification. Use periodic quality checks. A trained evaluator, supervisor, peer observer, or other designated person can periodically compare observations with those of regular data collectors. Repeated disagreement may signal that definitions, training, procedures, or measurement tools need to be reviewed. Determine how to review and adjust your observational system for the next evaluation Your observational system should itself be evaluated. Ask whether the system produced the information you needed, whether data collection was manageable, whether participants experienced the process as respectful, whether important perspectives or outcomes were missed, and whether any methods created unnecessary burden or privacy concerns. Use what you learn to improve the system for future evaluation cycles. With thoughtful planning, appropriate training, clear ethical safeguards, and ongoing review, an observational system can provide useful information for understanding and improving your work. In Summary To evaluate a program or community effort effectively, you need a systematic way to collect useful information about what is being implemented, how people experience the work, what conditions influence it, and what outcomes occur. An observational system provides that structure. It may include direct or participant observation, self-reports, administrative records, assessments, technology-assisted measures, or other methods depending on the evaluation questions. Observational systems should be feasible, participatory, ethically designed, and appropriate to the community and program context. Involving evaluators, data collectors, program staff, participants, and community members in design can improve the relevance of the information collected while helping identify concerns related to privacy, accessibility, interpretation, burden, and cultural context. Clear definitions, appropriate training, consistent procedures, and regular review can strengthen the reliability and usefulness of the resulting evaluation. Contributor Stephen B. Fawcett Phil Rabinowitz Resources Online Resource Assessing Children's Physical Activity in Their Homes: The Observational System for Recording Physical Activity in Children-Home, an article written by McIver et al. (2009), provides a real-life example of observational systems in evaluations. CDC Data Collection Methods for Program Evaluation: Observation is an article that helps users understand observation as a method for evaluation. The University of Wisconsin – Extension provides an article on Collecting Evaluation Data: Direct Observation. This article by Taylor-Powell and Steele extensively discuss what and how to observe in an evaluation. The National Collaboration on Childhood Obesity Research: Measures Registry is a searchable database of diet and physical activity measures relevant to childhood obesity research. The purpose of this registry is to promote the consistent use of common measures and research methods across childhood obesity prevention and research at the individual, community, and population levels. Obesity and public health researchers need standard measures to describe, monitor, and evaluate interventions, particularly policy and environmental interventions, and factors and outcomes at all levels of the socio-ecological model. Print Resources Bailey, J. (1977). A handbook of applied research methods in applied behavior analysis. Tallahassee, Florida: The author, Department of Psychology. (pp. 74-126). Fawcett, S.,et. al. (2008). Community Tool Box Curriculum Module 12: Evaluating the initiative. Center for Community Health and Development. University of Kansas.