# NoSQL, CAP and BASE — Database Fundamentals

Source: https://www.skillbyai.com/en/database-fundamentals/b-nosql

> Explain NoSQL database families and the CAP theorem.

## Different trade-offs for different needs

**NoSQL** ("not only SQL") databases relax parts of the relational model to gain horizontal scale, flexible schemas or specialised performance. Families: **key-value** (Redis, DynamoDB), **document** (MongoDB, Couchbase), **wide-column** (Cassandra, HBase, Bigtable) and **graph** (Neo4j). Many favour **denormalised** data shaped around queries and scale out by **partitioning** data across nodes. The **CAP theorem** states that when a **network partition** occurs, a distributed data store must choose between **consistency** (every read sees the latest write or returns an error) and **availability** (every request gets a non-error response). Systems are often described as **CP** (prefer consistency, such as HBase or etcd) or **AP** (prefer availability, such as Cassandra in its default settings), though many are tunable. **BASE** (Basically Available, Soft state, Eventual consistency) describes the AP style, in contrast to ACID. Today the lines blur: many NoSQL systems offer transactions, and distributed SQL databases ("NewSQL", such as Google Spanner, CockroachDB, YugabyteDB) provide ACID transactions at scale.

## Picking two during a partition

When the network splits, a distributed database chooses between consistency and availability.

![Two groups of server nodes separated by a jagged break line, with a balance scale above weighing a lock icon against an open door icon.](assets/figures/database-fundamentals/section-8-map.svg) — Figure 8.1 — The CAP trade-off during a network partition.

## Matching workloads to database families

A rough guide; many systems span several categories.

```text
workload                                         typical choice
-----------------------------------------------  ------------------------------
business transactions, reporting, joins          relational (PostgreSQL, MySQL)
session store, cache, counters, leaderboards     key-value (Redis)
product catalogue with varied nested attributes  document (MongoDB) or JSON columns
IoT/time-series writes at huge scale             wide-column (Cassandra) / time-series DB
friend-of-friend, fraud rings, recommendations   graph (Neo4j)
global ACID transactions at scale                distributed SQL (Spanner, CockroachDB)
```

## NoSQL is not schema-less

Data always has a schema; NoSQL just moves it from the database into application code. Without discipline, documents drift into many shapes that every reader must handle.

**Quiz:** According to the CAP theorem, what must a distributed database choose between during a network partition?

- [ ] Speed and cost
- [ ] SQL and NoSQL
- [x] Consistency and availability
- [ ] Indexes and joins

*Answer:* Consistency and availability. Under a partition, a system can either refuse some requests (C) or answer with possibly stale data (A).
