Data Model and Core Concepts
InfluxDB organizes data using a specific structure often referred to as Line Protocol. A single entry consists of a measurement, tags, fields, and a timestamp.
memory_usage,host=web01,zone=us-east value=72.5 1678900000000000000
This structure breaks down into several key components:
- Database: A logical container for data. Files are isolated per database on the storage disk.
- Retention Policy (RP): Defines how long data is kept. Each database has a default policy (usually
autogen), but custom policies can enforce expiry, such as retaining data for only 24 hours. - Measurement: Acts as a container for metrics, similar to a table in relational databases (e.g.,
memory_usage). - Tags: Key-value pairs used for metadata and indexing. Data is grouped by tag sets, which are sorted lexicographically.
- Fields: Key-value pairs representing the actual measurement data. Unlike tags, fields are not indexed.
- Timestamp: A precise timestamp (Unix nanosecond) associated with the data point, serving as the primary index within the TSM storage engine.
Comparison with Relational Databases
| InfluxDB Concept | Relational Database Equivalent |
|---|---|
| Database | Database |
| Measurement | Table |
| Point | Row |
Points, Series, and Shards
A Point represents a single discrete data record defined by a timestamp, a set of fields, and a set of tags.
A Series is a collection of points that share the same measurement, retention policy, and tag set. Data within the same series is stored sequentially on disk.
A Shard is a physical storage unit that holds data for a specific time range under a retention policy. For instance, one shard might hold data from 08:00 to 09:00, while the next holds data from 09:00 to 10:00. Each shard operates as an independent instance of the TSM storage engine.
Storage Engine Architecture
The Time-Structured Merge (TSM) engine consists of four primary components:
- Cache: Functions similarly to a MemTable in LSM trees. Incoming writes are stored here first. It has a configurable memory threshold (defaulting to 25MB). Once full, the cache is snapshotted and flushed to a TSM file.
- Write-Ahead Log (WAL): A transaction log ensuring durability. In the event of a system crash, the WAL allows the engine to recover recent writes not yet flushed to disk.
- TSM Files: The compressed, read-only data files stored on disk. Individual files are capped at 2GB.
- Compactor: A background process that merges smaller TSM files into larger ones to optimize storage efficiency and handle data deletion.
File System Structure
Data is persisted in three main directories:
meta: Contains metadata about the cluster state (meta.db).wal: Stores Write-Ahead Log files (.wal).data: Contains the actual TSM data files (.tsm), organized by database, retention policy, and shard ID.
Database Administration via CLI
InfluxDB supports interaction via the CLI, HTTP API, or client libraries. To access the command line interface with human-readable timestamps:
influx -precision rfc3339
Database Management
- List databases:
SHOW DATABASES - Create a database:
CREATE DATABASE system_monitor - Delete a database:
DROP DATABASE system_monitor - Select a database context:
USE system_monitor
Measurement Management
Measurements are created implicitly upon data insertion.
- List measurements:
SHOW MEASUREMENTS - Insert data (auto-creating measurement):
INSERT server_status,host=app01 uptime=4520,load=0.55Here,
server_statusis the measurement,hostis a tag, anduptime/loadare fields. - Delete a measurement:
DROP MEASUREMENT server_status
Retention Policies
Retention policies manage data lifecycle. InfluxDB does not support deleting individual records directly; instead, expired data is purged automatically based on these policies.
- View policies:
SHOW RETENTION POLICIES ON system_monitor - Create a policy (e.g., keep data for 7 days):
CREATE RETENTION POLICY "one_week" ON system_monitor DURATION 7d REPLICATION 1 DEFAULT - Modify a policy:
ALTER RETENTION POLICY "one_week" ON system_monitor DURATION 30d DEFAULT - Delete a policy:
DROP RETENTION POLICY "one_week" ON system_monitor
Continuous Queries
Continuous Queries (CQ) are automated queries that run periodically to process data, typically used for downsampling.
- Example: Downsample data every 30 minutes:
CREATE CONTINUOUS QUERY "cq_30m" ON system_monitor BEGIN SELECT mean(load) AS avg_load, max(load) AS peak_load INTO "one_week":downsampled_load FROM server_status GROUP BY time(30m), host ENDThis query calculates the average and peak load every 30 minutes and stores the results in a target measurement.
- List existing CQs:
SHOW CONTINUOUS QUERIES - Remove a CQ:
DROP CONTINUOUS QUERY "cq_30m" ON system_monitor