---
title: "Introduction to ortho_heter_endo_gmm in order to estimate heterogeneous peer effects when identity is orthogonal to eligibility"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Introduction to ortho_heter_endo_gmm in order to estimate heterogeneous peer effects when identity is orthogonal to eligibility}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---


```{r, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>"
)
```

```{r setup}
library(heterogeneouspeereffects)
```

# Heterogeneous Peer Effects

The package **"Heterogeneous Peer Effects"** aims to estimate individual responses within groups. Our contribution to the standard linear-in-means model is that we allow individuals to respond differently to the outcomes of their peers depending on both their **identity** and **eligibility** for treatment.

- **Identity** refers to an observable characteristic of an individual (e.g., race, gender).
- **Eligibility** indicates whether an individual is allowed to receive the treatment.

In the function `ortho_heter_endo_gmm`, we assume that identity and eligibility do not coincide, meaning that eligibility is orthogonal to identity. For instance, in the context of Progresa, we may assume that boys (resp. girls) are more influenced by the actions of their male (resp. female) peers than by their female (resp. male) peers. Gender, which is the relevant identity for social interactions, is orthogonal to being eligible for Progresa, since eligibility is only based on household income. Specifically, we distinguish between:

- **Within-group peer effects** (*theta within*): interactions between individuals sharing the same identity.
- **Between-group peer effects** (*theta between*): interactions between individuals of different identities.

Here, we estimate four parameters: 
- `theta_within` for identity 1 (e.g., male)
- `theta_within` for identity 2 (e.g., female)
- `theta_between` from male to female
- `theta_between` from female to male

We propose a simple methodology to identify and estimate the model using **partial population experiments**, where only a subset of individuals within a group is eligible for treatment, and the proportion of eligible individuals varies across groups. The estimation procedure relies on the **Generalized Method of Moments (GMM)**.

---

## Assumptions

Our method is based on the following key assumptions:

1. **Linear-in-Means Model with Panel Data**  
   - The outcome depends linearly on both the average outcome and the treatment status.

2. **Conditional Common Trends**  
   - In the absence of treatment, the average change in aggregate outcomes among eligible individuals in treated groups would have been the same as in control groups.

3. **Stable Share of Eligibles**  
   - The proportion of eligible individuals within a group remains constant over time.
   
4. **Randomized Experiment**  
   - Treatment is randomly assigned. In particular, groups receive the treatment independently of their share of eligible units for both identities and their share of identity (e.g., male) in the group.

5. **Common Support**  
   - For any observed share of males, eligible males, and eligible females in the reference population, there exist both treated and non-treated groups.

---

## Data Requirements

To apply this methodology, the data must meet the following criteria:

- Each group (e.g., districts) must contain both **Eligible** and **Non-Eligible** individuals and **Identity1** and **Identity2** individuals.
- Groups must be categorized as either **Treated** or **Non-Treated**.
- Example: In analyzing the impact of **Progresa**:
  - **Eligible male individuals** = Low-income family male students
  - **Eligible female individuals** = Low-income family female students
  - **Treated groups** = Households receiving the cash transfer program

The analysis is conducted at an **aggregate level**, considering average outcomes within each group.

---

## Estimation Procedure

To estimate the model, the following variables are required:

- `YM`: Average outcome for male individuals within a group (vector)
- `YF`: Average outcome for female individuals within a group (vector)
- `D`: Group binary treatment indicator (vector)
- `sM`: Share of male individuals in the group (vector)
- `sEM`: Share of eligible male individuals in the group (vector)
- `sEF`: Share of eligible female individuals in the group (vector)

---

## Usage Examples

### Estimating `delta`, `theta_within`, and `theta_between`:
```r
test_nocov <- ortho_heter_endo_gmm(YM, YF, D, sM, sEM, sEF)

# View results
print(test_nocov)
```
### Outcomes

This package return the estimated coefficients (and their p_values of ):
- direct effect of treatment (delta)
- intra group effect of treatment on male (theta_within_M)
- intra group effect of treatment on female (theta_within_F)
- inter group effect of treatment from male to female (theta_between_F_M)
- inter group effect of treatment from female to male (theta_between_M_F)

---

This package provides a flexible and robust framework for estimating heterogeneous peer effects in grouped data, making it applicable to various empirical settings, such as education, labor markets, and social networks.

