# Spark DataFrame

> Test in-memory Spark DataFrames in a pipeline (programmatic).

#  Spark DataFrame

Test Spark DataFrames in a pipeline before writing them to a data source. DataFrames are registered as named temporary views; multiple views are supported if the contract has multiple schemas. This connection is used programmatically from Python — no credentials are needed, the existing Spark session is reused.

## 1. Install

No extra is required (the Spark session comes from your environment):

```bash
uv tool install --python python3.11 --upgrade datacontract-cli
```

In a notebook or pipeline environment, install `datacontract-cli` as a library instead. See [Installation](../installation.md) for pip, pipx, and Docker.

## 2. Create a contract from your DataFrames

Inside an active Spark session, import the schema of registered tables or views:

```bash
datacontract import spark --tables my_table --output datacontract.yaml
```

The generated contract includes a `servers` entry of type `dataframe`:

```yaml
servers:
  - server: production
    type: dataframe
```

## 3. Test the DataFrame

Register the DataFrame as a temporary view named like the schema, then run the test with the Spark session:

```python
from datacontract.data_contract import DataContract

df.createOrReplaceTempView("my_table")

data_contract = DataContract(
  data_contract_file="datacontract.yaml",
  spark=spark,
)
run = data_contract.test()
assert run.result == "passed"
```

## 4. Let it catch a violation

The contract becomes valuable when it detects drift. Tighten an expectation — for example, mark a field as `required: true` or add a quality rule — and run the test again: `run.result` becomes `"failed"` and `run.checks` lists each violation, so the assert stops your pipeline before bad data is written.

## Reference

The Spark data type mappings: **[Spark DataFrame Reference](../reference/dataframe.md)**.
