> For the complete documentation index, see [llms.txt](https://lauradang.gitbook.io/notes/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://lauradang.gitbook.io/notes/machine-learning/research-papers/stackgan.md).

# Stack GAN

This paper's model architecture has many components, so I thought it would be good to layout the specifics of the architecture before implementing it.

## Model Architecture

![](https://868646840-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-LztfBhQUrZzyA7O_ZkJ%2Fuploads%2Fgit-blob-866f95d9d595f5dd0c4c595393330054f3fa8c0c%2Fstacked-gan.png?alt=media)

1. Stage-I GAN
2. Stage-2 GAN

## Stage-I GAN

**Input**: Text embedding of the text description $$(\varphi\_t)$$

### Conditioning Augmentation (CA)

**Purpose**: Create $$\hat{c\_0}$$ vector that captures the meaning of $$\varphi\_t$$ with variations.

**Process**: $$\varphi\_t$$ → FC layer → $$\mu\_0, \sigma\_0$$ → $$\mathcal{N}(\mu\_0(\varphi\_t),\sigma\_0(\varphi\_t))$$ → $$\hat{c\_0}$$ sampled from this Gaussian distribution

**Output**:
