Environment Configuration
Java Runtime Setup
Neo4j requires a compatible Java environment. For instance, the 3.5.x community edition necessitates JDK 11. After downloading and installing the JDK, ensure the JAVA_HOME environment variable is correctly configured to point to your installation directory.
Neo4j Deployment
Extract the Neo4j archive to a directory without special characters or spaces. Configure the NEO4J_HOME environment variable. Launch the database via the command line script. Once started, access the browser interface at http://localhost:7474. The default credentials are neo4j/neo4j, which must be changed upon first login.
Core Cypher Commands
Node Manipulation
Creating Nodes
The fundamental syntax for node creation:
CREATE (identifier:Label {key: value});
To instantiate a User entity:
CREATE (u:User {name: "Emma", score: 85 })
RETURN u;
If the node identifier isn't referenced later, it can be omitted:
CREATE (:User {name: "Liam", score: 85 })
Multiple nodes can be instantiated simultaneously:
CREATE (u1:User {name: "Emma", score: 85 }), (u2:User { name:"Olivia", score: 92 });
Nodes can also exist without labels:
CREATE (n {name: "Unlabeled"})
RETURN n
Retrieving Nodes
Basic retrieval syntax:
MATCH (identifier:Label)
WHERE identifier.key = value
RETURN identifier;
Finding all User instances:
MATCH (u: User) RETURN u;
Filtering by attribute directly:
MATCH (u: User{name:"Emma"}) RETURN u;
Applying a WHERE clause:
MATCH (u: User)
WHERE u.name="Emma"
RETURN u;
Updating Nodes
Modify attributes using the SET keyword:
MATCH (identifier:Label {key:value}) SET identifier.new_key = new_value;
Adjusting Emma's score:
MATCH (u: User{name:"Emma"}) SET u.score=100;
Multiple properties can be updated in a single statement. Neo4j allows adding new properties dynamically even if they didn't exist initially:
MATCH (u: User{name:"Emma"})
SET u.score=100, u.level='Expert';
Deleting Nodes
Locate the node with MATCH and remove it with DELETE:
MATCH (u: User{name:"Emma"}) DELETE u;
Attempting to delete a node with existing relationships will trigger an error. To remove nodes based on complex conditions:
MATCH (u: User) WHERE u.score>80 DELETE u;
Relationship Manipulation
Establishing Relationships
Relationships connect two nodes. Syntax for creation:
CREATE (id_1:Label)-[:RelType {key:value}]->(id_2:Label);
Linking Emma and Olivia with a Follows connection:
CREATE (u1:User{name: "Emma"})-[r:Follows{since:2019}]->(u2:User{name: "Olivia"})
Creating a relationship automatically generates the associated nodes if they don't exist. Python example using py2neo:
from py2neo import Graph
db = Graph("bolt://localhost:7687", auth=("neo4j", "password"))
def build_graph():
db.run("""CREATE (a:User {name: 'Alice'})-[:Follows {since: 2021}]->(b:User {name: 'Bob'})"""")
if __name__ == '__main__':
build_graph()
To link existing nodes without duplicating them, combine MATCH and MERGE (note the directional arrow):
MATCH (u1:User {name: "Emma"}), (u2:User {name: "Olivia"})
MERGE (u1)<-[r:Follows{since:2019}]-(u2);
Querying Relationships
Finding the connection between Emma and Olivia:
MATCH p=(:User{name: "Emma"})-[r:Follows]-(:User{name: "Olivia"}) RETURN p
Refining the search with WHERE:
MATCH p=(u1:User)-[r:Follows]-(u2:User)
WHERE u1.name="Emma"
AND u2.name="Olivia"
AND r.since=2019
RETURN p
Updating Relationships
Use SET to alter relationship properties:
MATCH p=(u1:User)-[r:Follows]-(u2:User)
WHERE u1.name="Emma"
AND u2.name="Olivia"
SET r.since=2022
RETURN p
Removing Relationships
MATCH p=(u1:User{name: "Emma"})-[r:Follows]-(u2:User{name: "Olivia"})
DELETE r
Caution: Executing DELETE p on a path variable will erase both the relationships and the nodes involved in that path.
If nodes within the path possess other relationships outside the matched pattern, DELETE p fails. To forcefully remove the path and all immediately adjacent relationships, utilize DETACH DELETE:
MATCH p=(n1:User{name:"Olivia"})-[r*1..2]-(n2) DETACH DELETE p
Path Discovery
Finding all paths between two entities. In production, restrict the depth to prevent timeout (e.g., r*1..3):
MATCH p=(n1:User{name: "UserA"})-[*]-(n2:User{name: "UserB"})
WITH reduce(s="", node IN nodes(p) | s + '->' + node.name) AS path
RETURN substring(path, 2, length(path))
Standard Operations
Eradicating Data (DELETE)
To remove a node and its attached relationships, use OPTIONAL MATCH:
MATCH (u:User{name: "Olivia"})
OPTIONAL MATCH (u)-[r]-()
DELETE u, r
Purging the entire database:
MATCH (n)
OPTIONAL MATCH (n)-[r]-()
DELETE n, r
Or the more concise:
MATCH (n)
DETACH DELETE n
Removing only isolated nodes:
MATCH (n)
DELETE n
Removing all relationships while preserving nodes:
MATCH (n)
OPTIONAL MATCH (n)-[r]-()
DELETE r
Removing Attributes and Labels (REMOVE)
Stripping an attribute:
MATCH (u:User{name: "Emma"})
REMOVE u.score
Stripping a label:
MATCH (n:User{name: "Liam"})
REMOVE n:User
RETURN n
Sorting (ORDER BY)
MATCH (p:User)
RETURN p
ORDER BY p.score DESC
Combining Results (UNION and UNION ALL)
UNION ALL retains duplicates:
MATCH (n:User {name: 'Emma'})-[:Colleague]->(f)
RETURN f.name AS name
UNION ALL
MATCH (n:User {name: 'Emma'})-[:Follows]->(f)
RETURN f.name AS name;
UNION deduplicates:
MATCH (n:User {name: 'Emma'})-[:Colleague]->(f)
RETURN f.name AS name
UNION
MATCH (n:User {name: 'Emma'})-[:Follows]->(f)
RETURN f.name AS name;
Pagination (SKIP and LIMIT)
MATCH (n)
RETURN n.property
SKIP 5
LIMIT 10
Iteration (FOREACH)
MATCH p=(start)-[*]->(finish)
WHERE start.name = 'A' AND finish.name = 'D'
FOREACH (n IN nodes(p) | SET n.visited = true)
Merging Data (MERGE)
MERGE acts as an UPSERT, creating the element only if it does not exist. Creating an identical node results in no change:
MERGE (p:User{name: "Liam", score: 50})
Using CREATE will always insert a duplicate, even with identical attributes (distinguished by internal ID).
Creating duplicate relationships with CREATE:
MATCH (u1:User {name: "Emma"}), (u2:User {name: "Olivia"})
CREATE (u1)-[r:Follows{since:2019}]->(u2);
MERGE ... ON CREATE sets properties only upon initial creation:
MERGE (n:User{name:"NewUser"})
ON CREATE SET n.createdAt=timestamp()
MERGE ... ON MATCH updates properties when an existing node is found:
MERGE (n:User{name:"NewUser"})
ON MATCH SET n.updatedAt=timestamp()
NULL Handling
Missing or undefined values are treated as NULL. Filtering for nodes lacking a property:
MATCH (p:User)
WHERE p.score IS NULL
RETURN p
IN Operator
MATCH (p:User)
WHERE p.score IN [85, 92]
RETURN p
Conditional Logic (CASE)
MATCH (n:User)
RETURN
CASE
WHEN n.name='Emma' THEN "Hello " + n.name
WHEN n.score>90 THEN "High score"
ELSE "Standard"
END AS result
Map Projections
RETURN {key: "value", list_key:[{inner:"map1"}, {inner:"map2"}]}
Pipeline with WITH
Passes intermediate results to subsequent query parts. Sorting before aggregation:
MATCH (n:User)
WITH n
ORDER BY n.name DESC LIMIT 3
RETURN collect(n.name)
Restricting branch expansion:
MATCH (n {name: "Liam"})--(m)
WITH m
ORDER BY m.name DESC LIMIT 1
MATCH (m)--(o)
RETURN o.name
Computing and updating attributes:
MATCH (n:User{name: "Liam"})-[:Follows]-(friend)
WITH n, count(friend) AS c
SET n.friendCount = c
RETURN n.friendCount
Expanding Lists (UNWIND)
Transforms a list into individual rows:
WITH [[10, 20], [30, 40], 50] AS nested
UNWIND nested AS x
RETURN x
Double unwinding:
WITH [[10, 20], [30, 40], 50] AS nested
UNWIND nested AS x
UNWIND x AS y
RETURN y
Deduplicating:
WITH [5, 5, 10, 10] AS coll
UNWIND coll AS x
WITH DISTINCT x
RETURN collect(x) AS set
Shortest Path Algorithms
Single shortest path:
MATCH p=shortestPath((n1:User{name: "Liam"})-[*1..2]- (n2:User{name: "Olivia"}) )
RETURN p
All shortest paths:
MATCH p=allShortestPaths((n1:User{name: "Liam"})-[*1..2]- (n2:User{name: "Olivia"}) )
RETURN p
String Matching
MATCH (p) WHERE p.name STARTS WITH "Em" RETURN p
MATCH (p) WHERE p.name ENDS WITH "iam" RETURN p
MATCH (p) WHERE p.name CONTAINS "liv" RETURN p
MATCH (n:User) WHERE n.name =~ ".*li.*" RETURN n
Logical Connectives
MATCH (p) WHERE size(p.name)>4 AND p.score>45 RETURN p
MATCH (p) WHERE size(p.name)>4 OR p.score>45 RETURN p
MATCH (p) WHERE p.name="Liam" XOR p.score<60 RETURN p
MATCH (p) WHERE NOT p.name="Liam" RETURN p
</code>
Indexes and Constraints
Creating an index:
CREATE INDEX ON :User(name)
Dropping an index:
DROP INDEX ON :User(name)
Unique constraint:
CREATE CONSTRAINT ON (p:User) ASSERT p.name IS UNIQUE
Existence constraint:
CREATE CONSTRAINT ON (p:User) ASSERT exists(p.name)
CREATE CONSTRAINT ON ()-[r:Follows]-() ASSERT exists(r.since)
Dropping constraints:
DROP CONSTRAINT ON (p:User) ASSERT p.name IS UNIQUE
Viewing schema:
CALL db.indexes();
CALL db.constraints();
Enforcing index usage:
MATCH (p:User{name: "Liam"})
USING INDEX p:User(name)
RETURN p
Label Exclusion
MATCH (n) WHERE NOT (n:Admin OR n:User)
Multi-Relationship Matching
MATCH p=(n1:User)-[r:Follows | :Colleague]-(n2)
RETURN p
Variable Depth Paths
MATCH p=(n1:User{name: "Liam"})-[r:Follows*1..3]-(n2)
RETURN p
Built-in Functions
Predicate Evaluation
exists()
Validates pattern or property presence:
MATCH (p:User)
WHERE exists(p.score)
RETURN p
MATCH (n)
WHERE exists(n.name)
RETURN n.name AS name, exists((n)-[:Follows]-()) AS has_connection
Metadata Extraction
MATCH (p:User{name: "Emma"}) RETURN keys(p), labels(p)
MATCH p=()-[]-() RETURN nodes(p), relationships(p)
Collection Validation
all()
Returns true if every element satisfies the condition:
MATCH (p:User{name: "Emma"}) SET p.vals = [1, 2, 3, 4, 5]
MATCH (p:User)
WHERE all(x IN p.vals WHERE x>0)
RETURN p
any()
Returns true if at least one element satisfies the condition:
MATCH p=()-[]-()
WHERE any(n IN nodes(p) WHERE n.score>40)
RETURN p
none()
Returns true if zero elements satisfy the condition:
MATCH p=()-[]-()
WHERE none(n IN nodes(p) WHERE n.score=40)
RETURN p
single()
Returns true if exactly one element satisfies the condition:
MATCH p=()-[r]-()
WHERE single(n IN nodes(p) WHERE n.score=85)
RETURN p
Scalar Functions
MATCH (n) RETURN id(n), properties(n)
Relationship Utilities
MATCH p=(n:User{name: "Emma"})-[r]-() RETURN type(r)
MATCH p=(n:User{name: "Emma"})-[r]-() RETURN startNode(r), endNode(r)
List Processing
MATCH (p:User{name: "Emma"}) RETURN head(p.vals), last(p.vals), size(p.vals)
coalesce()
Returns the first non-null value:
MATCH (p:User{name: "Emma"}) SET p.status=coalesce(null, "", "active")
extract()
MATCH p=()-[]-() RETURN extract(n IN nodes(p) | n.name) AS names
filter()
MATCH (p:User{name: "Emma"}) RETURN p.vals, filter(x IN p.vals WHERE x>3) AS filtered
Aggregation Functions
MATCH (p:User) RETURN count(p.score), avg(p.score), max(p.score), min(p.score), sum(p.score), collect(p.score)
String Manipulation
MATCH (p:User{name: "Emma"})
RETURN p.name, toUpper(p.name), lower(p.name), substring(p.name, 0, 2), replace(p.name, 'mm', "ll")
RETURN left("hello world", 5)
RETURN reverse("hello world")
RETURN trim(" hello ")
RETURN split("hello world", " ")
Sequence Generation
RETURN range(1, 10)
RETURN range(0, 10)[1]
RETURN range(0, 10)[1..3]
RETURN [x IN range(0, 10) WHERE x%2=0]
RETURN [x IN range(0, 10) | x^2]
CSV Data Import
Batch Node Creation
For a file with headers (users_h.csv):
name,score
Emma,85
Olivia,92
Liam,50
Move the file to the Neo4j import directory, then execute:
LOAD CSV WITH HEADERS FROM "file:///users_h.csv" AS row
MERGE (p:User{name: row.name, score: toInteger(row.score)})
For a headerless file (users.csv):
LOAD CSV FROM "file:///users.csv" AS row
MERGE (p:User{name: row[0], score: toInteger(row[1])})
Batch Relationship Creation
Given relations.csv:
source,target,duration
Emma,Olivia,2019
Emma,Liam,2020
Liam,Olivia,2021
Create connections:
LOAD CSV WITH HEADERS FROM "file:///relations.csv" AS row
MATCH (u1:User{name: row.source}), (u2:User{name: row.target})
MERGE p=(u1)-[r:Follows{since: toInteger(row.duration)}]->(u2)
Data Export via APOC
Install the APOC plugin by placing the corresponding JAR in the plugins folder and updating neo4j.conf:
apoc.export.file.enabled=true
apoc.import.file.enabled=true
dbms.security.procedures.unrestricted=apoc.*
Restart the service and verify with RETURN apoc.version(). Exported files land in the import folder:
WITH "MATCH (n:User) RETURN n.name AS name" AS query
CALL apoc.export.csv.query(query, "exported.csv", {})
YIELD file, source, format, nodes, relationships, properties, time, rows, batchSize, batches, done, data
RETURN file, source, format, nodes, relationships, properties, time, rows, batchSize, batches, done, data;
Python Integration (py2neo)
Entity Instantiation
from py2neo import Graph, Node, Relationship
db = Graph("bolt://localhost:7687", auth=("neo4j", "pwd"))
if __name__ == '__main__':
n1 = Node("User", name="Alice")
n2 = Node("User", name="Bob")
rel = Relationship(n1, "Follows", n2, since=2021)
db.create(rel)
Data Retrieval
from py2neo import Graph, NodeMatcher, RelationshipMatcher
db = Graph("bolt://localhost:7687", auth=("neo4j", "pwd"))
if __name__ == '__main__':
node_finder = NodeMatcher(db)
all_users = node_finder.match("User")
for user in all_users:
print(user)
target = node_finder.match("User", name="Alice").first()
print("Target: ", target)
rel_finder = RelationshipMatcher(db)
all_follows = rel_finder.match(r_type="Follows")
for link in all_follows:
print(link)
alice_links = rel_finder.match([target], r_type="Follows")
for link in alice_links:
print(link)
Attribute Modification
from py2neo import Graph, NodeMatcher, RelationshipMatcher
db = Graph("bolt://localhost:7687", auth=("neo4j", "pwd"))
if __name__ == '__main__':
node_finder = NodeMatcher(db)
target = node_finder.match("User", name="Alice").first()
target["score"] = 95
db.push(target)
rel_finder = RelationshipMatcher(db)
link = rel_finder.match([target]).first()
link["since"] = 2023
db.push(link)
Entity Deletion
from py2neo import Graph, NodeMatcher, RelationshipMatcher
db = Graph("bolt://localhost:7687", auth=("neo4j", "pwd"))
if __name__ == '__main__':
node_finder = NodeMatcher(db)
target = node_finder.match("User", name="Alice").first()
rel_finder = RelationshipMatcher(db)
link = rel_finder.match([target]).first()
db.delete(link)
db.delete(target)
Direct CQL Execution
from py2neo import Graph
db = Graph("bolt://localhost:7687", auth=("neo4j", "pwd"))
def execute_cql():
purge_cql = 'MATCH(n) OPTIONAL MATCH (n)-[r]-() DELETE n, r'
db.run(purge_cql)
build_nodes = 'CREATE(u1:User{name: "Emma", score: 85}), (u2:User{name: "Olivia", score: 92}), (u3:User{name: "Liam", score: 50})'
db.run(build_nodes)
fetch_all = "MATCH (u) RETURN u"
cursor = db.run(fetch_all)
for record in cursor:
print(record)
if __name__ == '__main__':
execute_cql()