# Data cleaning and lip calculations
Source: https://farmer-income-data-toolkit.org/conclusion/data-cleaning-and-lip-calculations
Once the full sample has been collected, data cleaning and calculation can begin. An R script template called “**Step 1: Data cleaning and calculation**” is available for guiding you through the necessary steps to arrive at the LIP analysis. The visual below gives an overview of the different steps taken in this data cleaning and calculation script.
**Living Income Price calculation**\
As seen in the overview above, the final part of the first R script is calculating the Living Income Price. Understandably, this is quite a crucial part of your calculations so let’s dive deeper into what that entails. The visual below provides an overview of the calculation.
The Living Income Price is essentially the yearly cost, divided by the quantity of produced crop. Yearly cost includes three key components:
**The Living Income Benchmark**, which is the cost of a decent but basic living standard for specific family size in a specific location. We adjust the benchmark to account for inflation and different family sizes. We also acknowledged that farmers might have multiple income sources, in which case the main crop should not be responsible for a 100% of the cost of living. This is why we multiply the benchmark by the diversification ratio, which is the proportion of income coming from the crop. Essentially, if a farmer earns 80% of their total income from coffee (diversification ratio = 0.8), then their income from coffee should be at least 80% of the Living Income benchmark. The other 20% should be earned from other sources the farmer has, such as livestock or the sale of other crops.
**The total production cost**, which is essentially the combination of all expenses directly related to the production of the crop. This is not accounted for in the Living Income benchmark and therefore needs to be included as an additional component to ensure that farmers cover their costs.
**The farm depreciation cost**, which is the total cost of establishing a farm divided over its expected lifespan to also account for initial farmer investments. This is a value, which you will have to decide on yourself as it varies from context to context. Work with your local partners to try and reach a reasonable value that corresponds to local conditions. In practice, if the cost of starting a coffee farm equals 1000 dollars and the farm can be productive for 30 years, then the farm depreciation cost is 1000/30 = **33.333.**
## **Data analysis and visualisation**
After data cleaning and calculation of all necessary variables comes the analysis, which you will find in the second R script called “**Step 2: Analysis and input for slide-deck**”. The script and presentation follow the same structure and include these sections:
| Section | Goal | Questions Answered |
| :--------------------- | :---------------------------------------------------- | :------------------------------------------------------------------------------------------------------------------------------------------------------ |
| 1. Sample | Understanding the farmers in your sample | How many farmers in the sample are women/youth/certified? Where are your farmers located? |
| 2. Income Sources | Assessing the role of the focus crop in farmer income | How much of farmer income comes from the crop? What other income sources do farmers have? What are other commonly grown crops? |
| 3. Production Cost | Getting insights into what makes up production cost | What are the most common cost drivers for farmers? What is the total cost of production? What types of labour do farmers have expenses for? |
| 4. Productivity | Comparing farmers based on production per acre | What is the median productivity per acre? How does this compare across different groups of farmers? |
| 5. Living Income Price | Exploring the LIP for different farmer groups | How high is the Living Income Price? How does it compare to the price farmers currently receive? Does the Living Income Price differ between groups? |
| 6. Interventions | Measuring impact of targeted interventions | By how much do interventions decrease the Living Income Price? |
**Tip: Use ChatGPT**\
Are pieces of the R script not working? Are you struggling to tailor the code to your dataset where variable names are slightly different? Asking ChatGPT or another AI Chatbot can be very useful to help you arrive at a working code without needing a lot of expertise.
**Final deliverable: Interactive slide deck**\
For sharing your analysis results, we recommend using Canva and Flourish. The combination of both platforms let you build an interactive slide deck, which means viewers can click on visual elements (for example, to switch currencies or drill down into specific breakdowns). Check out a live example of this interactive deck [here](https://www.canva.com/design/DAGrh9dXcSI/9cmn61SU7O_gt_YrRRNM0Q/view?utm_content=DAGrh9dXcSI\&utm_campaign=designshare\&utm_medium=link2\&utm_source=uniquelinks\&utlId=h445bc8c2d2).
To get started, make a copy of our “[LIP Presentation guidance](https://www.canva.com/design/DAGqCK1v2pc/I-sNxjJdPmVnjUchRUHWNQ/view?utm_content=DAGqCK1v2pc\&utm_campaign=designshare\&utm_medium=link\&utm_source=publishsharelink\&mode=preview)” template. It contains editable example slides alongside step-by-step instructions for creating and customising your visualisations.
## **How to arrive at recommendations from the analysis results?**
The data derived from the analysis can be used to develop targeted interventions to reduce or close the Living Income Gap. The interventions should be targeted to the specific context and there is no copy paste procedure here. The automatically generated heat-map at the end of the slide deck presents the scale of LIP reductions brought by different interventions. This is an option to ‘model’ the effect of specific interventions on the overall LIP.
Additionally to this, coming to recommendations requires a ‘manual’ approach. This includes discussion with the several stakeholders in your project about the interpretation of the data and how to use this data to inform decision making around targeted interventions in your specific context. There are several ways to look at developing interventions.
* ***Closing/reducing the LIP gap with price intervention:*** the difference between the median price and the median LIP provides an indication on the price intervention needed to close the living income gap. However, in many cases, the price gap is too big making it impossible to close the living income gap by price interventions alone
* ***Closing/reducing the LIP gap with targeted interventions:*** the data provides other insights that can help define targeted interventions to further reduce the living income gap. The most common ones are:
* ***Income diversification:*** In cases where farmers highly depend on the income coming from the main crop or where prices are relatively low it is worth investigating income diversifying activities that do not interfere with the (land used for) the main crop. Intercropping, off farm activities, livestock farming, agroforestry and carbon credit models are examples.
* ***Reduction of labour/production costs:*** the data defines the most common drivers for production costs. This can be labour activities, inputs (seeds / fertilisers), materials, non-mechanic equipment. If high production costs are driving up the living income price gap it is worthwhile to investigate targeted interventions to reduce costs in the most common cost drivers. Common examples are seedling projects, good (regenerative) agricultural practices (training and/or fertiliser programmes).
* ***Productivity:*** low yields are another common denominator of living income gaps. Increasing productivity very often reduces living income gaps if the interventions are well targeted and will receive a return on investment in the near future. Common examples here are farm management training, renovation of plants/trees, agroforestry system, soil management/soil fertility programmes.
* ***Group disaggregation:*** As farmers are not one homogenous group with the same needs and problems, disaggregation based on their characteristics allows us to pinpoint specific needs / interventions. Furthermore, interventions can be targeted based on the disaggregations provided in the analysis (e.g. gender, age, region, certification status).
## **Sensemaking and data validation**
Through organising local sensemaking sessions agency of data can be given back to the farmers that were part of the study. By validating findings with multiple stakeholders in the supply chains - first and foremost the farmers - accuracy, completeness and consistency of data can be checked, further enhancing data integrity and improving decision-making around targeting interventions.
Two models have been tested: a **workshop format** (structured, with slides and interactive exercises), and a **focus group format** (smaller, more informal discussions). The first not only validates data but also helps participants build capacity around concepts such as Living Income and co-create solutions; while the latter allows for deeper conversation and insights within a more intimate group.
Both served to return findings to farmers and other stakeholders, validate results and jointly explore implications for improving livelihoods.
**1) Sensemaking workshop**
Through a sensemaking workshop with a relevant sample of farmers, data can be validated and input for decision-making can be retrieved. When selecting the participants for the workshop, it is important to ensure each of the disaggregated groups researched (gender, age, certification status, region) are represented. A following agenda can be kept:
* **Introduction of the project:** a clear description of the project, the stakeholders, the objectives of the research and the sensemaking workshop
* **Introduction to the concept of Living Income:** to internalise the topic of Living Income a way of introducing it could be to ask participants to answer the following question: ‘what does a family need to have a decent life’? Followed by an internationally acknowledged definition of Living Income
* **Interactive exercises:** the discussion around the topic can be followed by a Living Income small-group exercise in which each group receives a given monthly income for the main crop and a budget sheet. The exercise can be done as follows:
* Cost of production exercise: The group calculates the cost of producing the main crop (e.g. labour, inputs, materials, non mechanic equipment) on a monthly basis.
* Expense exercise: The group makes an overview of the monthly household expenses (food, energy, housing, school, healthcare, others)
* Group discussion:
* Which household expenses can be met
* What expenses are prioritised
* Which expenses had to be left out
* What other sources of income are there, or which sources are needed
* Why is it important to understand the gap between income and decent income
* Bring back the discussion back to plenary and share learnings
* **Baseline result presentation:** a selection of the slide deck can be shared during the workshop. Make sure to use local language, valuta, metrics etc. Invite participants to respond to the results and verify if the outcomes of the research resonates.
* **Improving livelihoods:** follow the sharing of results with an interactive exercise (in smaller group) and organise discussions around the following questions:
* What are the barriers preventing you from having a good life?
* Can you prioritise the barriers or needs to improve livelihoods?
* What solutions or interventions could remove these barriers?
* What kind of support is needed?
* What successful interventions are already being implemented?
* **Reflexions and next steps:** conclude the session with summarising the main take outs of the workshop. Be clear on how the input will be used for future steps and / or setting up targeted interventions.
**2) Sensemaking focus groups**
When capacity and resources are short, this format will still enable rich, personal insights. Instead of a more formal workshop, the results of the research can be brought back in smaller (6-12 people), informal group discussions. The same agenda can be followed using printed material, allowing for validation of findings and gathering of insights in a more conversational way.
[^1]: Smallholders are defined by the producers, as the definition of smallholder varies per country and crop.
# Overview
Source: https://farmer-income-data-toolkit.org/conclusion/overview
**Links to tools:**
* R scripts → in Github
* [Slide deck manual](https://www.canva.com/design/DAGqCK1v2pc/I-sNxjJdPmVnjUchRUHWNQ/view?utm_content=DAGqCK1v2pc\&utm_campaign=designshare\&utm_medium=link\&utm_source=publishsharelink\&mode=preview)
# Farmer Income Data Toolkit: A practical methodology to assess living income gaps
Source: https://farmer-income-data-toolkit.org/introduction/farmer-income-data-toolkit
The Farmer Income Data Toolkit is the operational backbone of the **Living Income Commodity Strategy**. This technical, open-source toolkit helps you design, collect, analyse, and interpret farmer income data to define a living income price, the income floor farmers need for resilient and sustainable businesses, and ultimately, entire supply chains.
Developed by **Fairfood** and **Akvo**, this methodology is now fully available on GitHub, from survey design to analysis and recommendations, as well as case studies to inspire you while defining the outcomes of your project. We believe practical tools for fairer value distribution should be accessible to all, and this is a step towards more transparent, data-driven, and fair decision-making across value chains.
This toolkit is one of the legacies of the RECLAIM Sustainability!, a five-year program implemented by Fairfood, Solidaridad, Business Watch Indonesia, and Trust Africa in strategic partnership with the Dutch Ministry of Foreign Affairs. Learn more about the programme [here](https://fairfood.org/en/case/reclaim-sustainability-programme/).
# Introduction to the FID
Source: https://farmer-income-data-toolkit.org/introduction/introduction-to-the-FID
Across global agri-food supply chains, sustainability teams face growing pressure to move beyond ambition and demonstrate real impact: How are companies supporting equitable livelihoods in practice, not just in principle? New regulations such as the [EU Corporate Sustainability Due Diligence Directive (CSDDD)](https://commission.europa.eu/business-economy-euro/doing-business-eu/sustainability-due-diligence-responsible-business/corporate-sustainability-due-diligence_en) and the [EU Deforestation Regulation (EUDR)](https://environment.ec.europa.eu/topics/forests/deforestation/regulation-deforestation-free-products_en) are pushing businesses to take more responsibility for what happens at the start of their chains. This means finding concrete ways to support smallholder incomes, improve data quality, and strengthen long-term supply relationships.
The **Farmer Income Data Toolkit** helps you do just that.
Co-developed by **Fairfood** and **Akvo**, the toolkit is a core component of a **broader Living Income Commodity Strategy** jointly designed by **Fairfood, Akvo, and Heifer International**. This collaborative strategy aims to advance living income action in sectors where certification is limited and systemic inefficiencies persist. It equips supply chain actors with the evidence and approaches needed to make informed, fairer decisions. You can read more about the overarching strategy in the [**Living Income Commodity Strategy White Paper**.](https://fairfood.org/en/resources/heifer-and-fairfood-release-commodity-living-income-strategy-white-paper/)
This technical toolkit is your starting point for putting this strategy in practice by measuring farmer income and identifying the minimum income required for a viable, resilient livelihood. The methodology underpinning the toolkit builds on the [Income Measurement Guidance](https://idh.org/resources/income-measurement-guidance) developed by Akvo and IDH, with additional adaptations to support the analytical requirements of the Living Income Commodity Strategy mentioned above. It provides a step-by-step approach to gathering and analysing robust data, so you can move from insight to action with clarity and intention.
What sets this toolkit apart is its **segmentation-based approach**. Rather than treating smallholder communities as a single, uniform group, the toolkit enables you to uncover and explore key differences across farmer profiles. This added layer of analysis allows for **more nuanced and targeted interventions** that are grounded in real-world complexity, and tailored to performance, potential and context.
The toolkit is the culmination of a five-year programme funded by the Dutch Ministry of Foreign Affairs. The programme focused on testing and developing models that promote **fairer value distribution** as a foundation for resilient, inclusive, and ultimately more sustainable supply chains. In response, this Toolkit offers an end-to-end approach: from identifying income gaps to segment farmer realities, and support evidence-based decision making across supply chains.
### **A data-led approach to living income action**
At its core, our joint approach combines two complementary methodologies:
1. **Living Income Price (LIP).** Calculates a price floor that enables a decent standard of living for producers, whether at farmgate, cooperative, or Free-on-Board (FOB) level. By comparing current prices with a living income benchmark, it reveals how far the supply chain must move.
2. **Cost-Yield Efficiency (CYE).** Price alone rarely closes the gap, and that’s where the CYE assessment comes in. This analysis classifies farmers by both costs and yields, highlighting who is efficient, who is struggling, and why. This dual lens identifies where cost saving, productivity, or diversification measures might help - and where pricing interventions are really unavoidable.
Used together, these tools allow companies and their partners to identify income gaps with precision, and to design realistic interventions that are grounded in local data and developed in collaboration with producers. **Put simply, this is a practical roadmap for informed decision making and more effective investment in agri-food supply chains.** Based on traceable, locally sourced data, the approach outlines the steps in identifying and costing solutions for more equitable value distribution, shifting the conversation from compliance to collaboration. Instead of asking producers to *report*, it enables them to understand and *use* their data.
**This toolkit guides you through the first step** **of that journey.** It helps you calculate a fair and transparent price that supports a living income, and it also lays the groundwork for the next phase:developing practical, data-informed solutions through collective action with local stakeholders.
Whether you work in a **corporate sustainability team**, a **producer organisation**, an **NGO**, or a **certification body**, this toolkit is designed to help turn living income ambitions into **actionable, fundable, and implementable strategies**.
# Introduction to this toolkit
Source: https://farmer-income-data-toolkit.org/introduction/introduction-to-this-toolkit
### Who this guide is for
This guide is intended for organisations working in food supply chains, such as producer organisations, export organisation, NGOs, social enterprises, certification bodies, traders, or corporate sustainability teams, with the ambition of improving the living income of smallholder farmers. **It is particularly suited to users who are involved in impact measurement**, **value** **chain** **development**, or **price-setting initiatives**.
### What this guide is and what it is not
This guidance document outlines the methodology and steps required to carry out a Living Income Price (LIP) analysis. It supports users in developing a final slide deck that visualises the key findings of such an analysis. For each step in the process, we provide practical tools (including Excel templates, R scripts, and best practice documents) to help implement the methodology effectively. This toolkit is the result of a collaboration between Fairfood and Akvo, who have developed this lean version of the LIP analysis-approach based on field experience and iteration.
While the guide is highly detailed and comprehensive, it is not a fully automated tool. Users will need to manually complete various activities, make context-specific decisions, and adapt tools where necessary. From scoping to survey implementation, analysis, and reporting, a realistic timeline for carrying out the full LIP analysis is approximately two months. This guide is not meant to be a plug-and-play product or a black-box model - it is designed for users who want to understand and be actively involved in each step of the process.
The expected output of using this guide is a visually engaging slide deck that presents the key results of the LIP analysis and farmer income assessments. Insights include:
* a clear picture of farmer incomes, where they come from, and how they compare to the living income benchmark
* a deep dive into production costs, yields, and productivity trends
* focused disaggregations that spotlight different farmer groups and their unique realities
* evidence on how four targeted interventions can boost farmer incomes
* A price analysis revealing insights on the Living Income Price
A dummy version of the final slide deck can be found through [this link](https://www.canva.com/design/DAGrh9dXcSI/2vv1AffYwVKIFCvsHiqzUQ/view?utm_content=DAGrh9dXcSI\&utm_campaign=designshare\&utm_medium=link2\&utm_source=uniquelinks\&utlId=hc9ee034d3c).
### Whats needed for a successful LIP analysis
A successful LIP analysis depends on your ability to collect the required data. In some regions or value chains, farmers may not be used to tracking or reporting detailed cost and income information, which can affect the accuracy or completeness of your results. If you're not deeply familiar with the context, consider connecting with those who are. Local experts can help determine whether the necessary data can be collected and the correct approach to make farmers comfortable with the process. Local knowledge will always enhance the credibility and quality of the data. Lastly, it’s worth noting that the current version of the survey builder has been tailored for the following crops: banana, coffee, cocoa, spices, and shea. Other crops might need tailored questions to account for income correctly.
### Minimum skills required to use the toolkit
To use this toolkit effectively, users should possess the following baseline skills:
* **Survey and sample design**: Understanding how to frame relevant questions and select a representative sample of respondents is crucial.
* **Familiarity with indicator frameworks**: Users should understand basic concepts like household income, production costs, and farm profitability.
* **Basic R skills**: Data cleaning and analysis are performed using R scripts, so users should be able to run scripts, adjust parameters, and troubleshoot errors.
* **Data analysis literacy**: Users should understand foundational statistical concepts such as averages, medians, and standard deviations.
* **Excel proficiency**: The survey design tool is Excel-based and will be uploaded to KoboToolbox for data collection.
* Optional but helpful: comfort using **Canva** and **Flourish** to create engaging visuals in the final slide deck.
Note that this toolbox is modular. If certain steps fall outside yours or your team’ skillset , it’s always possible to outsource specific parts of the process. Both Akvo and Fairfood are available for support when needed. For technical support, data collection guidance, or help using the toolkit, contact: [info@akvo.org.](mailto:info@akvo.org) For insights on how to integrate this into your sustainability strategy or project design, contact: [info@fairfood.org](mailto:info@fairfood.org)
### Process Overview
The Living Income Price (LIP) analysis follows a clear three-phase process: **Scoping and preparation**, **Data collection**, and **Analysis and recommendations**. Each phase includes key activities that build toward a robust and context-specific LIP outcome.
During the **Scoping and preparation** phase, users define the goals, design the survey, and prepare for fieldwork. In the **Data collection** phase, enumerators are trained and fieldwork is conducted with ongoing data quality monitoring. Finally, the **Analysis and recommendations** phase involves cleaning and analysing the data, generating key insights, and visualising results in a final slide deck. This structured process ensures that the LIP analysis is both rigorous and practical, while remaining adaptable to different local contexts.
#### Contact for Support
This toolkit is designed to meet you where you are - whether you're just starting to measure income gaps or ready to co-create a pricing strategy with producers. Explore the tabs to get started, and remember: the process is modular, so you can go at your own pace, one step at a time.
* For technical support, data collection guidance, or help using the toolkit, contact: [info@akvo.org](mailto:info@akvo.org)
* For insights on how to integrate this into your sustainability strategy or project design, contact: [info@fairfood.org](mailto:info@fairfood.org)
# Resource Library
Source: https://farmer-income-data-toolkit.org/introduction/resource-library
As you navigate through the Toolkit, you may have questions about shaping your vision, and exploring the possibilities when defining your living income project. This Resource Library brings together materials that explain key concepts, showcase use cases from different contexts, and highlight practical approaches tested in the field.
Use them to inspire your project design, strengthen your rationale, and connect your interventions to proven strategies for achieving a living income.
## **White Paper:**
The Living Income Commodity Strategy builds on existing methodologies from Fairtrade International (Living Income Reference Price), True Price, and GIZ, while extending their applicability to farmers who are not certified. Within the repository, it explains the rationale behind the methodology and positions it as a practical mechanism to ensure that value concentrated at one end of the supply chain trickles down to farmers through targeted investments, clearly defined gaps to be closed, and income impact that can be monitored and quantified. This approach is an essential step in refining pricing and living income interventions: moving towards more transparent pricing mechanisms and genuinely sustainable value chains.
Access the document
Access the FAQ
## **Webinar: Learn from early adopters**
### 1. **Demo: The Farmer Income Data Toolkit**
This session provides a hands-on demonstration of how to use the Farmer Income Data Toolkit to generate insights and inform targeted interventions. Living income expert Lotje Kaak and data analyst Sandra Fudurova guided participants through the full process — from data collection and analysis to sensemaking — while answering key questions from the audience.
### 2. **Building a Living Income vision: Learn from early adopters**
In this webinar, hosted on **October 7th**, Andrea Moncada from **Molinos de Honduras–Volcafe** and Bless Agume from **Ndugu Farmers Cooperative (Uganda)** share how they are applying the Living Income Price Methodology in their own supply chains.
The session offers a first look at how farmer-level income data is being turned into actionable insights — from identifying income gaps to informing new interventions and measuring impact over time.
## **Case Studies repository:**
### 1. **Honduran Coffee**
Applied the Living Income Price (LIP) and Cost–Yield Efficiency (CYE) tools with Volcafe Honduras to map cost drivers, inefficiencies, and opportunities among supplier farmers.
Outcome: Influenced corporate policy through farmer segmentation as a crucial step in defining targeted interventions with clear KPIs and objectives, highlighting gender-based cost differences and efficiency gaps.
Use: The study validated results with 30 farmers during the [Laboratorio de Ingreso Digno](https://livingwagelab.org/all/first-in-country-living-income-lab-in-honduras/), fostering data ownership and integrating findings into the Volcafe Way, a farmer support programme to close income gaps.
Special feature: Direct private-sector integration, with recommendations embedded into ongoing supplier programmes.
**Use:** The study validated results with 30 farmers during the *Laboratorio de Ingreso Digno*, fostering data ownership and integrating findings into the “Volcafe Way” farmer support programme to close income gaps.\
**Special feature:** Direct private-sector integration, with recommendations embedded into ongoing supplier programmes.
Access it
### 2. **Sierra Leonean cocoa**
Following the creation of the first Living Income Benchmark by KIT in Sierra Leone, we applied the LIP and CYE in the country’s cocoa sector in collaboration with Solidaridad West Africa. Together, Fairfood and Solidaridad collected and analysed farmer-level data in five cocoa-producing districts, with Akvo supporting the identification of four farmer profiles.
Outcome: The first district-specific action points for input access, replanting, post-harvest handling, and youth inclusion.
**Goal:** Influence national cocoa policy. Findings will inform the Ministry of Agriculture and Produce Monitoring Board to engage the private sector on fair pricing.\
**Special feature:** National policy relevance, creating an evidence base for systemic interventions in an emerging sustainable cocoa market. The study was peer reviewed by the Produce Monitoring Board and the Sierra Leonean Ministry of Agriculture.
Access it
### 3. **Ugandan coffee**
Uganda became the third live use case of the Farmer Income Data Toolkit in 2025. Together with Ndugu Coffee Farmers and Wakuli, Fairfood and Akvo applied the Living Income Price (LIP) tool alongside a farmer segmentation process to understand income gaps among Robusta producers.
Outcome: The analysis enabled partners to identify which interventions are needed — and for which farmer groups — to close the living income gap. A new addition tested in Uganda was price modelling, used to show how income levels could fall short under different price fluctuation scenarios.
**Context:** Uganda has now surpassed Ethiopia as Africa’s largest coffee exporter, driven by record-high Robusta prices. Yet higher prices alone do not guarantee viable livelihoods, making data-led decision-making critical in a volatile market.\
**Special feature:** Forward-looking price modelling integrated into living income analysis, supporting farmer-first sourcing decisions that remain robust amid market volatility.\
**Use:** Findings are informing a new programme run by Fairfood, Wakuli and Ndugu with support from the Due Diligence Fund. The goal is to have data guiding discussions, sourcing strategies, and intervention design with the ambition to make Robusta coffee a viable and attractive livelihood and to demonstrate what farmer-first sourcing looks like in practice.
Access it
### 4. **Coming soon:**
The approach is being tested across different projects and programmes. This repository is expected to soon include findings from other projects, including:
* Uganda: Coffee (Wakuli), continuation of Ndugu via GIZ Due Diligence Fund (2025-2026)
* Uganda: Coffee (Ugacof), Developed within the EIT Food program (2025-2027)
* India: Spices, developed within the Social Sustainability Fund (SSF)
* Ghana: Shea, developed within the Social Sustainability Fund (SSF), together with Solidaridad West Africa
Register to be updated about new cases and events
# Data quality monitoring
Source: https://farmer-income-data-toolkit.org/process/data-collection/data-quality-monitoring
Data quality should be actively monitored during data collection. This is recommended because it allows you to do corrections on the data while the enumerators are still in the field.
KoboToolbox provides automatic visualisations that can be used to track the responses that come in and act quickly when an outlier is detected. You can do this tracking using the KoboToolbox feature of automatic charts through the “Data” menu, then choosing “Reports”. Here you will see visualisations of each variable collected.
The survey is very extensive so tracking every question will not be feasible. We recommend paying attention to the following variables, as they are most important to be of good quality for the LIP calculation:
* Quantity produced
* Quantity sold
* Measurement units
* Price
* Questions that will lead to labour cost
* Number of labourers
* Number of days for which labour is hired
* Salary per day
* Land size
* Costs: general farm costs, input costs, other costs
* Family size
**What to look out for?**\
Anything that seems very unusual for the local context. Talk to your local partners to determine what is possible and try to establish reasonable ranges for the key variables mentioned above. Some questions to answer together may include:
* What is the largest realistic quantity of crop that farmers can produce?
* Are there some local measurement units that are likely to be used by farmers?
* What is the maximum and minimum realistic price?
* What is the most a temporary labourer might get paid per day of labour?
Once you know what is realistic, you can start to flag values outside of these ranges. Remember, collected data sometimes does not confirm what we believe, so it is important to be cautious when removing or replacing values.
**Correcting the Data**\
Supervisors can directly contact enumerators to resolve any issues or to request clarification on unusual data entries. Ideally, the data gets corrected directly in KoboToolbox or the data collection app that is being used. Remember that we also collect phone numbers through the survey, so in some cases it is also possible to correct values over the phone.
When you are conducting a very large data collection, it might be worthwhile to connect your KoboToolbox database to a dashboard with your key variables of interest. An example of such a dashboard developed in Google Looker Studio can be found [here](https://lookerstudio.google.com/u/0/reporting/41ec1833-e98d-4533-bab5-cb72f2eb2ac0/page/p_mxm8sgy1od?s=qoj6eUWmhbo).
# Enumerator training
Source: https://farmer-income-data-toolkit.org/process/data-collection/enumerator-training
Once enumerators are identified, the data collector facilitates a one to two days training (depending on the number of the indicators to collect). The training workshop consists of learning how to use the data collection application on the smartphone (in the case of the farmer survey), understanding the survey, practising interviewing techniques, and learning how to troubleshoot during field data collection. Typically, a training session for the farmer survey data collection includes:
| Planning | Module | Description |
| :-------- | :------------------------ | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 0.5 day | Mobile data collection | Downloading mobile data collection application Getting to know all the features of the application Calibrating the GPS signal on the phone Battery saving options in the field |
| 0.5-1 day | Understanding the survey | Background and objectives of the project and the data collection Going through the survey question by question Survey best practices and informed consent Simulation in local languages |
| 0.5 day | Practicing (in the field) | Practicing the survey in the field or amongst each other (if the field is too far from the training location). Feedback from the practicing activity |
At the end of the training workshop, the enumerators have group discussions on challenges encountered in the field and have the opportunity to clarify questions about the survey.
[Click here](https://docs.google.com/document/d/1coYfKYc2H5e6PfgrmyhM0E8AbxjK1uU0znHPmRM33xk/copy) to make a copy of the enumerator field guide, that contains a prepared informed consent and other necessary guidance regarding the survey questions.
# Ethical considerations
Source: https://farmer-income-data-toolkit.org/process/data-collection/ethical-considerations
Various ethical principles are considered during the data collection phase.
**Informed consent and participation**
Data collectors adhere to the principle of 'do no harm.' Every participant is fully informed about how their data will be used and is asked for consent through a specific question built into the survey. This ensures that participation is voluntary and informed.
**Data accessibility and security**
* ***Access control***: Access to survey data is restricted to specific staff members involved in the data collection or analysing the survey. The raw data will be anonymised before doing any analysis and no personal information will be used for analysis.
* ***Non-disclosure***: The raw data is kept confidential and cannot be shared with external stakeholders without a data sharing agreement.
**Data storage and handling**
* ***Cloud storage***: Data is initially stored in its original format on the KoboToolbox server.
* ***Data cleaning and access***: The data cleaning and access process should take into account your organisation’s internal data practices and policies.
# Overview
Source: https://farmer-income-data-toolkit.org/process/data-collection/overview
**Links to tools:**
* [Enumerator field guide](https://docs.google.com/document/d/1coYfKYc2H5e6PfgrmyhM0E8AbxjK1uU0znHPmRM33xk/copy)
# Procedures and tips for a quality interview
Source: https://farmer-income-data-toolkit.org/process/data-collection/procedures-and-tips-for-a-quality-interview
This section provides instructions on maximising the chances of obtaining the respondent's cooperation and conducting effective interviews.
**Before any interview**\
Before conducting fieldwork in any community, the field team should contact the local authorities and inform them about the data collection. Even if the list of respondents is available, it is important to notify the local authorities and relevant stakeholders before starting any interviews.
**Obtaining informed consent from respondents**\
It is the responsibility of each enumerator to contact respondents and persuade them to participate in the study. Obtaining respondents' consent can sometimes be challenging, especially in the case the respondents have been pre-identified. The success of the study depends on the ability of enumerators and team leaders to convince selected respondents to participate.
Here are some tips to maximise the chances of obtaining their consent to participate:
* **Dress professionally and always wear your identification badge.** A good presentation is essential to inspire trust in your respondent.
* **Make the respondent feel comfortable.** Start the interview with a smile and a polite greeting before introducing yourself, rather than jumping straight into the questionnaire.
* **Obtain verbal informed consent before collecting any data.** A consent text explains the purpose of the study, clarifies that participation is entirely voluntary, ensures that responses are anonymous, and informs respondents that they have the right to refuse to answer any question or stop the interview at any time if they feel uncomfortable.
* **Maintain a confident and positive approach.** Avoid apologetic phrasing or questions that invite refusal, such as *"Are you too busy?"* Instead, say, *"I’d like to ask you a few questions"* or *"I’d like to speak with you for a few moments."*
* **Show the official authorisation for data collection if necessary.**
* **Emphasise anonymity when needed.** If the respondent is hesitant or asks about the purpose of the data, explain that all collected information will remain strictly anonymous, no names will be used, and the final report will anonymise all responses.
* **Answer all respondent questions honestly.** Before agreeing to an interview, the respondent may have questions about the survey, their selection, or other concerns. Be direct and pleasant in your responses. If they are worried about the interview duration, provide an estimated time without exaggerating, as this might discourage them.
* **Do not make false promises or offer incentives** to persuade respondents to participate.
* **If the respondent is unavailable, offer to return at a more convenient time.** Agree on a specific date and time for the interview.
* **If the respondent refuses outright, politely ask (without insisting) if they are willing to share their reasons for refusal,** as this information needs to be reported.
* **At the end of the interview, thank the respondent** for their time and participation.
**The consent form**\
Before asking any questions, the interviewer is required to present the contents of the consent form to the respondent. The consent form is structured around the following points:
1. Introduction of the interviewer and their employer. As mentioned earlier, a friendly introduction will help put the respondent at ease and facilitate the interview process
2. Explanation of the study’s objectives.
3. Notification of the interview duration.
4. Presentation of confidentiality protocols.
5. Information on the voluntary and risk-free nature of participation in the study.
6. Respondent’s questions.
7. Useful contacts.
**Behaviour during the interview**
Once the respondent's consent has been obtained, the interview can begin. The following principles must be followed throughout the process:
**Interview the respondent alone:** The presence of another person during the interview may prevent you from obtaining honest responses. Therefore, it is crucial that the interview is conducted in private. If others are present, explain to the respondent that their answers are confidential and that it is necessary to conduct the interview in a setting conducive to one-on-one conversation. In all cases where others are present, make every effort to isolate yourself with the respondent as much as possible. You may, for example, suggest finding a comfortable place to sit in order to answer the questions, which can also serve as a reason to move away from noisy children or other distractions that could disrupt the interview.
**Being neutral during the interview:** Most people are polite and tend to give answers they believe you want to hear. Therefore, it is crucial that you remain completely neutral when asking questions. Never allow your facial expressions or the tone of your voice to make the respondent think they have given the "right" or "wrong" answer. Never appear to approve or disapprove of the respondent's answers.
All questions are carefully formulated to maintain neutrality. They do not suggest that one answer is more likely or preferable than another. If you do not ask the full question, you may compromise this neutrality. You can remind the respondent, if necessary, that there are no right or wrong answers.
If the respondent gives an ambiguous answer, try to probe while maintaining neutrality by asking questions such as:
* “Could you elaborate a bit more?”
* “I didn’t quite understand. Could you repeat, please?”
* “There is no urgency. Take your time to think it over.”
If the respondent does not understand the question, it is best to first repeat the question exactly as it is formulated in the questionnaire. Often, a simple rereading of the question can resolve the misunderstanding.
Silence can be an important tool for probing the respondent. If a respondent provides an incomplete or inadequate answer, remaining silent for a few seconds after their response often encourages them to elaborate or clarify their answer.
The interviewer must always remain attentive to any difficulties the respondent may encounter, as well as to their emotions and any signs of frustration. For example, if the interviewer notices that the respondent is unsettled for any reason and is no longer in the best condition to answer the questions, it may be helpful to suggest a short break.
**Never suggest answers to the respondent:** If the respondent's answer to a question is not relevant, do not prompt them by saying something like, "*I suppose you mean... don't you?*" In many cases, the respondent will agree with your interpretation of their answer, even if that is not what they actually meant. Instead, you should probe further so that the respondent provides the relevant answers themselves.
**Never change the wording or sequence of the questions:** The wording and sequence of the questions in the questionnaire must be maintained. If the respondent does not understand a question, you should repeat it slowly and clearly. If there are still issues, you can read it again, being careful not to change its original meaning. Only provide the minimum required information to obtain an appropriate answer.
**Do not try to explain the questions:** It is important to standardise the administration of the questionnaire as much as possible by all interviewers. Therefore, it is essential that each interviewer limits themselves to reading each question to the respondent exactly as written. If each interviewer starts explaining the questions, they will end up having different meanings for different surveys, and the responses will be less useful for future analysis. Explain only if the questionnaire instructs you to do so. Otherwise, if a respondent asks what a question means, ask them to answer based on their own understanding.
**Do not attempt to define the words used in the questions:** Unless there are instructions (or definitions) in the questionnaire, leave it to the respondent to interpret each word according to their understanding. Key terms will be explained during training.
**Pace and tone of voice:** If the interviewer reads the questions with a monotonous or hesitant tone, or if the pace of the interview is not appropriate (too fast or too slow), the respondent may lose interest in the questions, become less cooperative, and provide poor-quality information. The interviewer must therefore know how to modulate the pace and tone of their voice depending on the respondent, their reactions, and their attitude.
**Managing hesitant respondents tactfully:** There will be situations where respondents may simply say "I don't know," give an inappropriate answer, appear bored or detached, or contradict something they’ve already said. In these cases, you need to try to re-engage them in the conversation. For example, if you feel they are shy or scared, try to alleviate their shyness or fear before asking the next question. Take some time to talk about things unrelated to the interview (e.g., their neighborhood/village, the rainy season, their daily activities, etc.).
Respondents should not be allowed to go off on tangents or provide unnecessary details for the survey. This is because there is a risk that the later questions will often be rushed. Therefore, time management for the interview is essential. When dealing with very "scattered" and talkative individuals, kindly steer their responses back to the questions asked. Try to bring them back to the core topics while still being attentive.
As an example, if a respondent gives irrelevant or complicated answers, don’t interrupt them abruptly or rudely, but listen to what they have to say. Then, gently guide them back to the original question. A good atmosphere must be maintained throughout the interview.
The best atmosphere for an interview is one where the respondent views the interviewer as a friendly, compassionate, and sensitive person who does not intimidate them and to whom they can say anything without feeling shy or embarrassed. As mentioned earlier, the main issue in gaining the respondent's trust is often discretion. This issue can be prevented if you are able to secure a private location for the interview.
If the respondent is reluctant or unwilling to answer a question, explain once again that the same question is being asked of other household heads (both women and men) in the same environment, and that the responses will be grouped together. If the respondent remains reluctant, simply write "refusal" (or the corresponding code) for that question and proceed as if nothing happened. Remember, the respondent cannot be forced to provide an answer.
**Don’t rush the interview:** Ask questions slowly to ensure that the respondent understands what is being asked. After asking a question, pause and give the respondent time to think. If the respondent feels rushed or is not allowed to formulate their opinion, they may respond with "I don’t know" or give an inaccurate answer.
If you feel the respondent is answering without thinking just to speed up the interview, tell them there is no rush to answer, and their opinion is very important, so they should consider their answers carefully.
That’s why it’s crucial to read the text in the questionnaire word for word, or its closest possible translation, without omitting anything or adding anything. The translations of key terms provided during training should be used. If the person responds before you have finished reading the question, continue reading and then ask them to confirm their answer (especially when listing proposed answers).
Stay focused on your questionnaire and avoid distractions so you don’t waste the respondent’s time or the interview. Stick to the questionnaire’s allotted time.
**Don’t leave a question without an answer:** Unless the respondent clearly says they don’t want to answer, never leave a question with the idea of returning to it later. As mentioned earlier, questions should be asked in the order they appear in the questionnaire, while following the instructions related to skips.
**Takeaway – The good interviewer:** Asks the questions exactly as they are written in the questionnaire, in the same order, ensuring that the respondent feels comfortable and is able to understand and answer.
# Field preparations
Source: https://farmer-income-data-toolkit.org/process/scoping-and-preparation/field-preparations
Field preparation consists of three main steps:
The selection of enumerators is crucial and is based on the required data points and the available time for data collection. Based on the survey length and the context (e.g., travel and field conditions), the local project team can decide on the number of interviews to be carried out per day. The formula below is used to calculate the final number of enumerators, with the number of days allocated for data collection and the number of data points per enumerator adjustable as needed.
Enumerators are selected based on the following criteria:
* Background in agriculture
* Understanding of basic financial concepts, such as profit and loss
* Fluency in the local language or dialect
* Familiarity with the geographic area
**Tip: Use a Terms of Reference (ToR) in nearby universities**\
Enumerators who have previously worked with the company are always preferred for their familiarity with the area and local farmers. If new enumerators are needed, good practice is to circulate a Terms of Reference (ToR) to nearby universities with agricultural programmes to ensure the recruitment of knowledgeable candidates. An example of such a ToR can be found [here](https://docs.google.com/document/d/1Eo5Zc4ENu8tPu862W0GjrCAI1idEXAIirN7sf_TCGvc/edit?tab=t.0):
After enumerators are selected, a detailed data collection plan is formulated. This plan specifies which enumerators will visit particular farmers and on which days. The data collection should be done primarily by the project team. Finalising this plan during the kickoff ensures preparedness and coordination.
The following considerations should be considered while designing the data collection plan:
* **Travel time**: Ensure travel time is evenly distributed among enumerators.
* **Flexibility**: Include extra farmers each day to compensate for potential non-attendance.
* **Communication**: Farmers should be informed about their interviews at least one day in advance.
* **Support:** Organise enumerators into groups for mutual support and equip them with field guides to assist in locating farmers.
# Implement survey into KoboToolbox
Source: https://farmer-income-data-toolkit.org/process/scoping-and-preparation/implement-survey-into-KoboToolbox
Follow these steps to upload the adjusted survey from the Google Sheets into KoboToolbox.
**Step 1: Download the Survey Builder as an Excel file**
Open the file directly in Excel. A few practical adjustments are needed to ensure the survey looks correct in the KoboToolbox app.
* **“Survey” tab:** Review the "Disabled" column. Any row marked as TRUE will be excluded from the survey. If you wish to keep a question, change the cell value to FALSE. Check orange cells, which have been adapted to include the name of your focus crop.
* **“Choices” tab:** Check the orange cells. These multiple-choice options were automatically adjusted based on the information in the Contextualisation Sheet. Perform a final review to ensure all required options are included.
* The **contextualisation sheet** does not need to be removed.
**Step 2: Upload the file into KoboToolbox**
* If you do not have a KoboToolbox account yet, go to [this web page](https://eu.kobotoolbox.org/accounts/login/) to create one for free
* Add “New project”
* Click on "Upload an XLSForm" and select the downloaded survey builder file from your local files.
* Fill in the project details: Complete the required project details, such as the project name and description, then proceed with the upload.
If you encounter any errors during the upload process, use the official [KoboToolbox XLSFrom validator](https://getodk.org/xlsform/) to identify and resolve the issue. Read the error message carefully and make the necessary corrections in the downloaded Excel file.
**Step 3: Review the survey in Kobo**
* Preview the survey: Once the survey is successfully uploaded, go to the "Preview form" section to review how it appears within KoboToolbox.
* Edit the survey (if necessary): If you need to make further adjustments or add additional questions, you can click on "Edit form" to make changes. Another option is to adjust the questions in the survey builder excel file and reupload to Kobo using the same procedure as described above.
**Collecting data with KoboCollect app**\
You have now successfully uploaded your survey to KoboToolbox! The next step is to make sure the enumerators can access the survey on their KoboCollect application. To have access to the survey on their phone, the enumerators should follow the steps below (using internet connection):
**Step 1: Install KoboCollect**
* Install KoboCollect via the Play Store or Google Play (There is no KoboCollect for IOS users)
* Open the application and click on either “QR code” or “manually enter project details.”
* If manual: Fill in the 3 fields displayed as follows:
* URL ==> [https://eu.kobotoolbox.org](https://eu.kobotoolbox.org/) (or [https://kf.kobotoolbox.org](https://kf.kobotoolbox.org/) )
* “Username”⇒ XXXX
* “Password” ⇒ YYYYWeD\@kobo2023
* Click on “Add”
**Step 2: Add a project on KoboCollect**\
If KoboCollect is already installed and you want to add a new project (account), click on the upper right corner of the home screen of the application (see image below) and select “Add project” to provide the account information for the new project
**Step 3: Download a form/survey**\
Before you can fill out a form/survey, you first need to have the form/ survey on the KoboCollect app on your phone. To download the form/survey, you need to go to:
1. The home screen of the KoboCollect app, click on “download form”
2. Select the form(s) you are interested in, and then click on “Get selected.”
Now that you have access to the form/survey on the KoboCollect app on your phone, you can proceed with filling out the survey offline and other tasks. The major tasks to be done by the enumerators are:
* **Task 1: Fill out a form/survey:** To fill out a form/survey, the enumerator should click on “start new form” on the home screen of the KoboCollect application and select the form/survey to be filled out. This action should be repeated each time the enumerator wants to conduct an interview with a given form/survey.
* **Task 2: Save or finalise the form/survey:** A form/survey for a given interview can be either saved partially or finalised.
* ***Save as draft***: A form can be saved as a draft during the interview by pressing the save button (disk icon, see image below) at any question in the form/survey, or at the end of the survey, click on “save as draft” (see image below).
* ***Finalise a form***: Make sure all the required questions are filled in. At the end of the form/survey, you will see “You are at the end of \[Survey Name]. Click on “Finalise” to end the survey.
Save as draft
Finalise a form
* **Task 3: Send a finalised form**: Once a form/survey is finalised, it moves to the sub-menu “ready to send”. To send the finalised form/survey, the enumerator need to :
* First, make sure you have access to internet connection.
* On the KoboCollect home screen, click on the sub-menu “*ready to send*”
* Select the interviews you want to send to the server
* And click on send
# Intake and kick off
Source: https://farmer-income-data-toolkit.org/process/scoping-and-preparation/intake-and-kick-off
The initial step of the scoping and preparation phase involves a kickoff meeting where the data collection team gathers essential information using **the contextualisation form** (you can find this sheet in the first tab of the [survey builder](https://docs.google.com/spreadsheets/d/1_9KPzlAJ-X6oedozcBpwATZTlL5ZLW-4/copy)). This meeting aims to:
* Introduce the different parties involved in the Living Income Price (LIP) analysis.
* Present the survey content, discussing specific data needs for the particular context with the local partner.
* Review and fill out the contextualisation form.
The contextualisation form is designed to efficiently gather the required information to properly tailor the survey to the local context and avoid missing any critical context details for field data collection. It serves several key purposes:
* Defines the scope criteria (see table below).
* Adjusts the survey to reflect local contexts, such as customising multiple-choice answers to enhance data quality and reduce outliers.
* Provides contextual validation for data, offering insights into reasonable ranges for specific variables like farmgate prices.
The responsibility to fill out the contextualisation form ideally sits with the local partner that actively works with the farmers and knows the context. If needed, a local agronomic expert can be involved to ensure survey relevance.
## Scope Criteria of Data Collection
| **Factor** | **Scope** |
| ------------------------- | -------------------------------------------------------------------------------------------------------------------------------- |
| **Population definition** | Smallholder farmers\*\* |
| **Sample size** | See sampling section later in this chapter (uses population-based approach) |
| **Geography** | Define the specific areas (e.g. provinces, districts…) |
| **Focus crop** | 1 focus crop |
| **Reference period** | The last 12 months → Important to properly calculate the LIP and compare with the Living Income benchmark (also based on 1 year) |
\*\*Smallholders are defined by the producers, as the definition of smallholder varies per country and crop.
### Filling Out the Contextualisation Form
The **contextualisation form** has two main parts:
**1. Survey context**, which transforms a generic survey template to a context specific survey, with the correct multiple choice options and both questions and structure adjusted to the focus crop. The goal here is to essentially customise the choices for multiple-choice questions, to best reflect what farmers may commonly answer and improve the relevance of the responses. With most questions, farmers also have the option to answer “Other” after which they will be prompted to specify what they meant. This means that it is not necessary to input every single possible answer, but rather the most common ones, which are likely to be answered by multiple farmers. The contextualisation sheet further includes hints and prompts to provide further guidance.
**2. Parameter value** part, which will be crucial during data quality monitoring and analysis. The Living Income Price analysis includes a lot of variables, which makes data quality essential in deriving correct results. This part of the contextualisation form is meant to ease the process by setting up an expectation about common ranges in the data.
The contextualisation sheet also includes a manual checklist for tasks that cannot be automated.
# Overview
Source: https://farmer-income-data-toolkit.org/process/scoping-and-preparation/overview
**Links to the tools:**
* [Survey builder](https://docs.google.com/spreadsheets/d/1_9KPzlAJ-X6oedozcBpwATZTlL5ZLW-4/copy)
* [Enumerator ToR](https://docs.google.com/document/d/1Eo5Zc4ENu8tPu862W0GjrCAI1idEXAIirN7sf_TCGvc/edit?tab=t.0)
# Sampling
Source: https://farmer-income-data-toolkit.org/process/scoping-and-preparation/sampling
The Farmer Survey data collection utilises a **population-based sampling approach** to ensure that the sample accurately represents the larger population of farmers involved in the projects. This method calculates the required sample size based on the total population size and a predetermined significance level, both of which are essential for maintaining statistical validity and reliability in the survey results.
**Key elements of the sampling approach:**
1. ***Sample size calculation***: The sample size is determined using a formula that takes into account the total population size and the desired confidence level (typically 95%) and margin of error (commonly set at 5%). This ensures that the sample is large enough to produce statistically significant and generalisable results. For example, a larger population requires a larger sample to maintain the same level of confidence and precision in the results.
2. ***Population frame***: To determine the appropriate sample size, the M\&E local team needs a comprehensive list of all the farmers involved in the project. This list serves as the foundation for defining the population frame—the group of individuals from which the sample will be drawn. The population frame should be as complete and accurate as possible, encompassing all relevant farmers within the scope of the project. If there are multiple project locations or distinct groups within the farmer population (e.g., different farming practices, crops, or regions), the sampling approach may need to be stratified to ensure that these variations are adequately represented.
3. ***Adjustments for non-response***: When calculating the sample size, it’s also important to account for potential non-responses or incomplete data. Typically, a buffer (e.g., increasing the sample size by 10-20%) is included to compensate for this possibility and maintain the desired level of statistical power.
In the case that the population frame is higher than 300 and the sample size needs to be calculated, several reliable online tools are available. A recommended resource is [Raosoft's sample size calculator](http://www.raosoft.com/samplesize.html).
In scenarios where a comprehensive farmer list is unavailable or cannot be obtained, alternative sampling strategies such as the snowballing or referral method are employed.
* ***Snowballing technique:*** This method is useful in reaching populations that are difficult to sample when a list is not available. It starts with a small group of known respondents (seeds) who then refer to other potential respondents within their network, growing the sample size akin to a rolling snowball.
* ***Referral method:*** Similar to snowballing, this technique relies on initial respondents to refer subsequent participants. This method is particularly valuable when targeting specific characteristics within a population that are known only to insiders (e.g., certain types of farmers or farming practices).
**Tip: Don't forget to note plans for disaggregation**
When analysing data, it’s crucial to disaggregate by variables such as gender or region to ensure representative results. Properly allocating the sample size across these disaggregation groups, like targeting an adequate number of farmers of each gender, is essential. To facilitate this, preferences for disaggregation should be identified early, and necessary variables must be included in the farmer list. Insufficient initial information can lead to a sample that is unsuitable for disaggregation, potentially resulting in biased outcomes if analysis proceeds under these conditions.
# Survey design
Source: https://farmer-income-data-toolkit.org/process/scoping-and-preparation/survey-design
The LIP Producer Survey is a quantitative tool developed to systematically capture the indicators needed for a Living Income Price analysis. This survey is designed for in-person interviews, allowing direct engagement with respondents. This format facilitates the collection of accurate and detailed information. The survey’s structure is flexible and tailored to meet the specific needs of a LIP study while adaptable to the contexts of different crops and countries.
### Content of the survey
The LIP Producer Survey is composed of 10 sections, each focusing on key areas relevant to the LIP analysis:
1. **Introduction and consent:** Questions to introduce the survey and obtain consent from participants.
2. **Household demographics:** Questions on the farmer’s background, including age, gender, education, and household composition. This section is primarily used for data disaggregation.
3. **Farm characteristics**: Information about the farm, such as its size, location, and the variety of the focus crop.
4. **Revenue from the focus crop**: Questions on the quantity produced and sold, price, types of buyers, and revenue from the previous season.
5. **Production cost**: Information on input costs, asset costs, labor, etc.
6. **Income from other crops:** Questions on the quantity produced and sold, prices from other crops and the different types of other crops.
7. **Income from livestock:** Questions on the types of livestock, the revenue from livestock, the related labour and other costs.
8. **Income from off-farm labour and other income:** Assess all other income sources and their related income.
9. **Access to finance and related costs**: Information about access to finance, including sources of loans.
10. **Cooperative Membership:** Questions about membership in agricultural cooperatives, its benefits, and certification.
The primary goal of this survey is to accurately measure the Living Income Price (LIP). The details of how different questions contribute to this LIP are outlined in the calculation sheet.
### How to use the Survey Builder
A Survey Builder tool has been developed to facilitate the design of the survey, ensuring consistency across different cases while allowing customisation to specific contexts. To begin, open the survey builder by making a copy of the template through [this link](https://docs.google.com/spreadsheets/d/1_9KPzlAJ-X6oedozcBpwATZTlL5ZLW-4/copy).
The survey layout is designed for direct upload into **KoBoToolbox**, a widely recognised platform for digital survey development. This format offers flexibility and adaptability, ensuring that the survey can be tailored to the unique needs of each context.
The workbook consists of the following tabs:
* Contextualisation sheet
* Survey (for KoboToolbox upload)
* Choices (for KoboToolbox upload)
* Settings (for KoboToolbox upload)
### KoboToolbox Survey, choices, and settings templates
The **Survey Worksheet** contains the full list of questions and their details, defining how they will appear and function in the KoBoCollect app or other compatible data collection platforms. This sheet should not be modified, as it is automatically contextualised based on the selections made in the Contextualisation Sheet. The main columns of the survey worksheet are:
| Column name | Explanation |
| :----------------- | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Type** | This column specifies the type of input expected for each question. Examples include text, integer, decimal, select\_one, select\_multiple, date, and more. The type defines how the question will be presented in the app and determines the format of the data collected. |
| **Name** | No two entries can have the same name. Names have to start with a letter or an underscore. Names can only contain letters, digits, hyphens, underscores, and periods. Names are case-sensitive. |
| **Label** | The column contains the actual text you see for the question. |
| **Hint** | Provides additional guidance to enumerators or respondents on how to answer the question. Hints can clarify the intent of the question or provide examples to help respondents understand what is being asked. |
| **Disabled** | This column determines whether a question will be displayed or hidden. If set to TRUE, the question is disabled and will not appear during data collection. In this context, the Disabled column is linked to the Information Worksheet, where indicators and preferred methods are selected. Questions related to unselected indicators are automatically disabled (set to TRUE). |
| **Appearance** | |
| **Default** | |
| **Relevant** | This column controls the visibility of a question based on conditions or responses to previous questions. For example, if a respondent selects "Yes" for a specific question, additional follow-up questions can be triggered. Conversely, irrelevant questions can be skipped. |
| **Constraint** | Allows you to impose restrictions on the respondent’s answers to ensure data quality. For instance, a constraint can enforce that a respondent enters an age greater than 18 or that a numeric value falls within a specified range. When a respondent provides an invalid response, an error message is displayed, guiding them to correct their input. |
| **Calculation** | It is used to perform calculations based on the values of preceding questions. The result is stored in a calculated variable, which can then be utilised in question labels or to set constraints for subsequent questions. |
| **Repeat\_count** | Instead of allowing an infinite number of repeat group questions, this column can be used to specify the exact number of times a set of questions should repeat. |
| **Choice\_filter** | This column enables responses to a question to filter the options available in subsequent questions by implementing a cascading select. |
The **Choices Worksheet** defines the options for multiple-choice questions. Each row represents a single choice, and these are grouped by **list\_name**, linking them to the relevant question in the Survey Worksheet.
It contains 3 main columns:
| Column name | Explanation |
| :------------- | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **list\_name** | Specifies the name of the choice list associated with a question. This name links the choices to a specific question in the Survey Worksheet, ensuring the correct options are displayed. Example: For a question about crop types, the list\_name could be crop\_list, which ties the options (e.g., maize, rice, wheat) to the relevant question. |
| **name** | Serves as the unique identifier for each choice. This is the value that will be stored in the dataset when a respondent selects an option. We commonly use numbers from 1. Additionally we use: **0** for options such as “None of the above” **9999** for “I don’t know” **9998** for “I prefer not to say” **77** for “Other” or “Other, please specify:” |
| **label** | Displays the text or description of the answer choice as seen by the respondent. |
While optional, the **Settings Worksheet** is highly recommended to ensure proper metadata management for the survey. At a minimum, you should specify the **form\_title** and **form\_id**, which are essential for organising and identifying your survey. The **form\_id** defaults to the date, but it can be customised if needed.
For more information on using KoBoToolbox’s XLSForm format, please refer to [https://xlsform.org/en/](https://xlsform.org/en/)
### Manual adjustments
After making a copy of the survey and filling out the contextualisation sheet, there are a few leftover tasks that cannot be automated. These can be found under the **manual checklist** part of the contextualisation form and are mainly related to the **choices** sheet. The lists that require attention are in red.
You will need to manually fill in the names of the enumerators that are being sent to the field and the location names for every administrative level. It is only necessary to fill in the names of localities where the data collection is taking place as these will be used to record the location of the farmer. Make sure that the list\_name is **exactly the same** for all options and that the name column contains only **unique** **values**.
Example: Data collection in Sierra Leone in 3 districts.
| list\_name | name | label |
| :--------- | :--- | :--------- |
| admin2 | 0 | Not Listed |
| admin2 | 1 | Port Loko |
| admin2 | 2 | Kambia |
| admin2 | 3 | Bombali |
The second adjustment is related to the crop harvests and seasons. When we ask farmers about their production quantities and cost, we allow them to choose the time period for which they want to report. In practice, some farmers may choose to report their quantities and cost for the main and the off season separately. Some may have only had one productive season. Some may know their cost values but only per the whole year. The survey is currently designed to let farmers answer for the last 12 months, season 1 or season 2. That is determined by this list in the **choices** sheet:
| list\_name | name | label |
| :------------------------ | :--- | :-------------------------- |
| f\_focus\_rev\_timeperiod | 1 | Whole year (past 12 months) |
| f\_focus\_rev\_timeperiod | 2 | Season 1\* |
| f\_focus\_rev\_timeperiod | 3 | Season 2\* |
However, that might not be optimal for every focus crop. In some scenarios the seasons also have distinct local names. For example, in the context of coffee in Uganda this could be the Main crop and the Fly Crop. Labels with an \* can be replaced by these local names to facilitate better understanding between enumerators and respondents. If you think farmers might prefer reporting by four seasons instead of two feel free to adjust the list:
| list\_name | name | label |
| :------------------------ | :--- | :-------------------------- |
| f\_focus\_rev\_timeperiod | 1 | Whole year (past 12 months) |
| f\_focus\_rev\_timeperiod | 2 | Season 1\* |
| f\_focus\_rev\_timeperiod | 3 | Season 2\* |
| f\_focus\_rev\_timeperiod | 4 | Season 3\* |
| f\_focus\_rev\_timeperiod | 5 | Season 4\* |
\*to be replaced by more distinctive/local names\
It is important to choose the option that will feel the most natural to farmers to get more accurate and reliable information.
Lastly, the **manual checklist** also includes a list of tasks for the contextualisation sheet that ensure the automatic adjustments to the survey work as intended. In case you do not require the Spanish translation of the survey, delete the following columns from the **survey** sheet:
* label::Spanish, hint::Spanish, constraint message::Spanish