Introduction
nitomic is Datomic's peer API (datomic.api) ported to
clonim, a Clojure compiler that
targets Nim. Datomic programs written against datomic.api compile to native
binaries, with no JVM involved.
(require '[datomic.api :as d])
(d/create-database "datomic:mem://hello")
(def conn (d/connect "datomic:mem://hello"))
@(d/transact conn [{:db/ident :person/name
:db/valueType :db.type/string
:db/cardinality :db.cardinality/one}])
@(d/transact conn [{:person/name "Ada"} {:person/name "Grace"}])
(d/q '[:find [?n ...] :where [_ :person/name ?n]] (d/db conn))
;=> ["Ada" "Grace"]
A clean-room port, checked against the real thing
Datomic Pro 1.0.7705 ships its peer library as AOT-compiled JVM classes, with no Clojure source. nitomic is therefore a clean-room reimplementation of the peer API in Clojure that clonim can compile.
It is checked against Datomic itself. The same test programs run on the JVM against Datomic Pro and natively against nitomic, and their outputs must match line for line. The main one of those programs is Datomic's own getting-started walkthrough, which is what the Getting started part of this book explains step by step.
The bootstrap database is also taken from Datomic. The system partitions,
value types and attributes, with Datomic's own entity ids, docs and
transaction ids, were dumped from a fresh datomic:mem database. New entity
ids are allocated the way Datomic allocates them:
- transaction ids are
tin partition 3; - other ids share the
tcounter in their partition; - attributes get their own counter in the db partition.
So the ids a program sees are the ids Datomic would give it.
How to read this book
- Installation and Your first program get a program compiled and running.
- Getting started walks through the Seattle
example in
examples/seattle/, explaining each step and showing the output it produces. - The Reference part lists what is supported, where nitomic differs from Datomic, and how the port is tested.
Installation
nitomic is Clojure source, not a compiled library. You need
clonim to compile programs, and you
point clonim at nitomic's src/ directory with --source-path.
Requirements
- Nim 2.2.12 or later, with Nimble.
- clonim, at a revision from
mainthat includes codegod100/clonim#1. That change added the Clojure features nitomic relies on:reify/deftype/defprotocol,#inst/#uuidand*data-readers*,compareand comparator sorts,for/doseqmodifiers,ex-data, and a set of collection functions. - On Linux, the PCRE runtime library (
libpcre3on Debian and Ubuntu).
If you don't have Nim yet, choosenim is the quickest route:
curl https://nim-lang.org/choosenim/init.sh -sSf | sh
export PATH="$HOME/.nimble/bin:$PATH"
With Nimble
nitomic is a Nimble package, and installing it also installs clonim:
nimble install https://github.com/codegod100/nitomic
Nimble installs src/ as is (the package keeps only .clj files), so the
package directory is itself the source path:
clonim run hello.clj --source-path "$(nimble path nitomic)"
From a checkout
git clone https://github.com/codegod100/nitomic
cd nitomic
nimble install # installs the working tree (and clonim)
A checkout also gives you two Nimble tasks:
| task | what it does |
|---|---|
nimble example | runs the Seattle walkthrough, examples/seattle/getting_started.clj |
nimble test | runs every test program and diffs its output against test/expected/ |
Both use the clonim on your PATH. Set CLONIM to use a different one:
CLONIM=path/to/clonim/bin/clonim nimble example
Building clonim by hand
This is what CI does, and it is handy when you want a specific clonim revision:
git clone https://github.com/codegod100/clonim
cd clonim
nim c --hints:off --warnings:off -o:bin/clonim src/clonim.nim
Then use path/to/clonim/bin/clonim wherever this book says clonim.
Your first program
Save this as hello.clj:
(require '[datomic.api :as d])
(d/create-database "datomic:mem://hello")
(def conn (d/connect "datomic:mem://hello"))
;; Install one attribute.
@(d/transact conn [{:db/ident :person/name
:db/valueType :db.type/string
:db/cardinality :db.cardinality/one}])
;; Add two people.
@(d/transact conn [{:person/name "Ada"} {:person/name "Grace"}])
(prn (sort (d/q '[:find [?n ...] :where [_ :person/name ?n]] (d/db conn))))
Run it
clonim run hello.clj --source-path path/to/nitomic/src
clonim run compiles the program through Nim and runs it. It prints:
("Ada" "Grace")
The query returns its results in no particular order, so the program sorts them before printing.
Build a binary
clonim build hello.clj --source-path path/to/nitomic/src
clonim build produces a standalone native executable instead of running
the program. Release builds (-d:release) are what you want for speed: built
this way, the full Seattle walkthrough runs in about 1.3 seconds.
What just happened
The program uses nothing but datomic.api, so the same file also runs on the
JVM against Datomic Pro. A few things to notice:
datomic:mem://hellonames an in-memory database, gone when the program exits.datomic:sql://hello?jdbc:sqlite:hello.dbwould keep it in a SQLite file instead (see Durable storage).d/transactreturns a future. Dereferencing it with@waits for the result, a map with:db-before,:db-after,:tx-dataand:tempids. A failed transaction throws when you dereference it.- Maps without
:db/idget implicit tempids. Each map in the second transaction becomes a new entity. [?n ...]is a collection find spec: the query returns a vector of names instead of a set of tuples.
The Getting started walkthrough covers all of this, and much more, on a realistic data set.
Overview
Datomic Pro ships a getting-started tutorial built around a small data set
about Seattle's neighborhood communities: blogs, mailing lists, Twitter
feeds, chambers of commerce and so on. The tutorial lives in the Datomic
distribution as samples/seattle/getting-started.clj, a file of forms meant
to be evaluated one at a time in a REPL.
nitomic carries that tutorial as a single program:
examples/seattle/
├── getting_started.clj # the walkthrough, as one program
├── seattle-schema.edn # the schema: 10 attributes and their enums
├── seattle-data0.edn # the initial data: 150 communities
└── seattle-data1.edn # more data, added later: 108 more communities
The data files come from the Datomic Pro distribution, which is licensed under the Apache License 2.0.
Running it
From the root of a nitomic checkout:
nimble example
# or, equivalently
clonim run examples/seattle/getting_started.clj --source-path src
The program reads the .edn files by relative path, so run it from the
repository root.
Why it prints the way it does
The program runs unchanged on the JVM against Datomic Pro and natively against nitomic, and the two outputs are compared line by line (see How it is tested). Query results are sets, whose iteration order differs between the two implementations, and instants differ from run to run. So instead of printing results directly, the program prints them through two small helpers that produce a canonical form:
(defn- canon-str [x]
(cond
(map? x) (str "{" (str/join ", " (sort (map (fn [[k v]] (str (canon-str k) " " (canon-str v))) x))) "}")
(set? x) (str "#{" (str/join " " (sort (map canon-str x))) "}")
(vector? x) (str "[" (str/join " " (map canon-str x)) "]")
(seq? x) (str "(" (str/join " " (map canon-str x)) ")")
(inst? x) "#inst"
:else (pr-str x)))
(defn show
"Print a labelled result; an unordered collection of results is sorted."
[label x]
(println (str label ":") (canon-str x)))
(defn show-sorted [label xs]
(println (str label ":") (str "(" (str/join " " (sort (map canon-str xs))) ")")))
showprints a label and a value. Map keys and set members are sorted, and every instant prints as#inst.show-sortedprints a collection of results as a sorted list.
Every output line quoted in the following chapters is taken from
test/expected/getting_started.out, which was recorded on the JVM against
Datomic Pro 1.0.7705. nitomic must reproduce it exactly.
The chapters
| chapter | covers |
|---|---|
| The Seattle data model | the schema: communities, neighborhoods, districts, enums |
| Creating a database and loading data | create-database, connect, #db/id tempids, transact |
| Entities and pull | entity, navigation and reverse navigation, pull in queries |
| Queries | find specs, joins, :in parameters, predicates, fulltext |
| Rules | named, reusable, composable query clauses |
| Time travel | as-of, since, with |
| Changing data | partitions, adding and retracting values, retractEntity, the tx report queue |
The Seattle data model
The schema in examples/seattle/seattle-schema.edn describes three kinds of
entity, linked by references:
community ──:community/neighborhood──▶ neighborhood ──:neighborhood/district──▶ district
│ │
├─ :community/type ─▶ enum (:community.type/...) └─ :district/region ─▶ enum (:region/...)
└─ :community/orgtype ─▶ enum (:community.orgtype/...)
Attributes
| attribute | type | cardinality | notes |
|---|---|---|---|
:community/name | string | one | fulltext |
:community/url | string | one | |
:community/neighborhood | ref | one | a neighborhood |
:community/category | string | many | fulltext |
:community/orgtype | ref | one | an orgtype enum |
:community/type | ref | many | type enums |
:neighborhood/name | string | one | :db.unique/identity |
:neighborhood/district | ref | one | a district |
:district/name | string | one | :db.unique/identity |
:district/region | ref | one | a region enum |
Each attribute is installed with a plain map:
{:db/ident :community/name
:db/valueType :db.type/string
:db/cardinality :db.cardinality/one
:db/fulltext true
:db/doc "A community's name"}
Datomic (and nitomic) install an attribute implicitly when a transaction
asserts :db/ident, :db/valueType and :db/cardinality for a new entity;
no :db.install/_attribute is needed.
Unique identities
:neighborhood/name and :district/name are :db.unique/identity. When a
transaction asserts one of those values for a tempid, and an entity already
has that value, the tempid resolves to the existing entity instead of
creating a new one. This is called upsert. It is what lets the second batch
of data (seattle-data1.edn) mention "Beacon Hill" again without creating a
second Beacon Hill.
Fulltext
:community/name and :community/category are :db/fulltext, which makes
them searchable with the fulltext query function (see
Queries).
Enums
Enumerated values are entities with nothing but a :db/ident. Refs to them
can be written as the keyword:
[:db/add #db/id[:db.part/user] :db/ident :community.type/twitter]
| enum | values |
|---|---|
:community.orgtype/... | community commercial nonprofit personal |
:community.type/... | email-list twitter facebook-page blog website wiki myspace ning |
:region/... | n ne e se s sw w nw |
In an entity, a ref to an enum reads back as its keyword. In a query, you
join through :db/ident to get the keyword, or pass the keyword as an
input.
The data
The data files are vectors of maps. Each map carries a #db/id tempid, and
references between new entities use the same tempid:
{:district/region :region/e, :db/id #db/id[:db.part/user -1000001], :district/name "East"}
{:db/id #db/id[:db.part/user -1000002], :neighborhood/name "Capitol Hill",
:neighborhood/district #db/id[:db.part/user -1000001]}
{:community/category ["15th avenue residents"],
:community/orgtype :community.orgtype/community,
:community/type :community.type/email-list,
:db/id #db/id[:db.part/user -1000003],
:community/name "15th Ave Community",
:community/url "http://groups.yahoo.com/group/15thAve_Community/",
:community/neighborhood #db/id[:db.part/user -1000002]}
seattle-data0.ednholds 150 communities with their neighborhoods and districts.seattle-data1.ednholds 108 more, added in Time travel.
Creating a database and loading data
Requiring the API
(require '[datomic.api :as d]
'[datomic.db]
'[clojure.string :as str])
datomic.api is the whole public API. datomic.db provides id-literal,
the reader function behind #db/id[...] tempid literals, which the data
files use.
Creating and connecting
(def uri "datomic:mem://seattle")
(show "create-database" (d/create-database uri))
(def conn (d/connect uri))
create-database: true
create-database returns true when it creates a database, and false if
one with that name already exists. connect returns a connection, which is
what you transact against and take database values from.
nitomic:
datomic:memand every other protocol exceptsqlname an in-memory database in the same process. To keep a database on disk, use a SQLite URI such asdatomic:sql://seattle?jdbc:sqlite:seattle.db, or ajdbc:postgresql:one to share it between machines (see Durable storage).
Reading EDN with tempid literals
The schema and data files contain #db/id[:db.part/user] and
#db/id[:db.part/user -1000001] forms. To read them, bind
*data-readers* so the db/id tag goes to datomic.db/id-literal:
(defn read-edn [path]
(binding [*data-readers* {'db/id datomic.db/id-literal}]
(read-string (slurp path))))
Each literal becomes a tempid in the named partition. Two literals with the same negative number are the same tempid, which is how the data files link new entities to each other. A literal without a number is a fresh tempid.
Transacting the schema
(def schema-tx (read-edn "examples/seattle/seattle-schema.edn"))
(show "first schema statement" (first schema-tx))
(def schema-report @(d/transact conn schema-tx))
(show "schema tx-data count" (count (:tx-data schema-report)))
first schema statement: {:db/cardinality :db.cardinality/one, :db/doc "A community's name", :db/fulltext true, :db/ident :community/name, :db/valueType :db.type/string}
schema tx-data count: 75
d/transact returns a future. Dereferencing it (@) waits for the
transaction and returns a report map:
| key | value |
|---|---|
:db-before | the database value before the transaction |
:db-after | the database value after it |
:tx-data | the datoms the transaction asserted or retracted |
:tempids | a map from tempids to the entity ids they resolved to |
The 75 datoms are the 44 attribute definition datoms, one
:db.install/attribute datom per attribute (10), the 20 enum idents, and
the transaction's own :db/txInstant. nitomic allocates the same entity ids as
Datomic, so the same 75 datoms come out.
If the transaction fails, dereferencing throws an ex-info whose data
carries a Datomic :db/error code, such as :db.error/unique-conflict.
Transacting the data
(def data-tx (read-edn "examples/seattle/seattle-data0.edn"))
(show "first data statement" (dissoc (first data-tx) :db/id))
(show "second data statement" (dissoc (second data-tx) :db/id :neighborhood/district))
(def data-report @(d/transact conn data-tx))
(show "data tx-data count" (count (:tx-data data-report)))
first data statement: {:district/name "East", :district/region :region/e}
second data statement: {:neighborhood/name "Capitol Hill"}
data tx-data count: 1237
The :db/id keys are dropped before printing because a tempid prints
differently on the two platforms.
Some things this one transaction exercises:
- Map form. Each map asserts all its attributes for one entity.
- Tempid references.
:neighborhood/district #db/id[:db.part/user -1000001]points at the district created earlier in the same transaction. - Enum keywords as ref values.
:district/region :region/eresolves the keyword to the enum entity. - Cardinality many.
:community/category ["events" "news"]asserts one datom per value. - Upsert. A neighborhood that appears twice under two tempids resolves to
one entity, because
:neighborhood/nameis a unique identity.
With the schema and data in place, the database holds 150 communities.
Entities and pull
There are two ways to get at the attributes of an entity: the entity API, which gives a lazy, navigable map-like object, and pull, which gives plain data in the shape you ask for.
Finding the communities
(def results (d/q '[:find ?c :where [?c :community/name]] (d/db conn)))
(show "communities" (count results))
communities: 150
(d/db conn) returns the current database value, an immutable snapshot.
The query finds every entity ?c that has a :community/name; the value
position of the pattern is left out, so it matches any value. The result is
a set of one-element tuples.
Entities
(def id (ffirst (sort-by first results)))
(def entity (-> conn d/db (d/entity id)))
(show "entity keys" (set (keys entity)))
(show "entity name" (:community/name entity))
entity keys: #{:community/category :community/name :community/neighborhood :community/orgtype :community/type :community/url}
entity name: "15th Ave Community"
d/entity returns an entity for an id. Attributes are fetched lazily when
you look them up with a keyword or get. keys lists the attributes the
entity has.
The walkthrough sorts the results and takes the smallest id, rather than
using ffirst directly, so that both platforms pick the same community.
Navigating references
A ref attribute returns another entity, so you can walk the graph:
(let [db (d/db conn)]
(show-sorted "names and neighborhoods"
(map #(let [entity (d/entity db (first %))]
[(:community/name entity)
(-> entity :community/neighborhood :neighborhood/name)])
results)))
names and neighborhoods: (["15th Ave Community" "Capitol Hill"] ["Admiral Neighborhood Association" "Admiral (West Seattle)"] ...)
Reverse navigation
Prefixing the attribute name with _ follows a reference backwards. From a
neighborhood, :community/_neighborhood returns every community that points
at it, as a set of entities:
(def community (d/entity (d/db conn) (ffirst (sort-by first results))))
(def neighborhood (:community/neighborhood community))
(def communities (:community/_neighborhood neighborhood))
(show-sorted "communities in the same neighborhood" (map :community/name communities))
communities in the same neighborhood: ("15th Ave Community" "CHS Capitol Hill Seattle Blog" "Capitol Hill Community Council" "Capitol Hill Housing" "Capitol Hill Triangle" "KOMO Communities - Captol Hill")
Other entity behaviour worth knowing:
- a cardinality-many attribute returns a set;
- a ref to an enum returns the enum's keyword;
d/touchloads every attribute, and a touched entity prints as its attribute map;- in nitomic, an untouched entity prints as
{:db/id n}.
Pull in a query
Put a pull expression in :find to get maps instead of ids:
(def pull-results (d/q '[:find (pull ?c [*]) :where [?c :community/name]] (d/db conn)))
(show "pull results" (count pull-results))
pull results: 150
The pattern [*] pulls every attribute. Refs come back as nested maps
holding :db/id, and cardinality-many attributes come back as vectors. Here
is the belltown community with its refs reduced to their keys:
a pulled community: {:community/category ["events" "news"], :community/name "belltown", :community/neighborhood (:db/id), :community/orgtype (:db/id), :community/type ((:db/id)), :community/url "http://www.belltownpeople.com/"}
A pull pattern can name just the attributes you want, and a :find can mix
pulls with plain variables:
(d/q '[:find ?n (pull ?c [:community/url])
:where [?c :community/name ?n]]
(d/db conn))
names with urls: (["15th Ave Community" {:community/url "http://groups.yahoo.com/group/15thAve_Community/"}] ...)
Outside queries, d/pull and d/pull-many take a database, a pattern and
an entity id (or ids). Patterns support nested maps for refs, reverse
attributes, recursion, and the :as, :limit and :default options.
Queries
d/q takes a query and its inputs, and the first input is usually a
database. A query is data: a vector (or a map, or a string) with :find,
optional :in and :with, and :where.
(d/q '[:find ?c :where [?c :community/name]] (d/db conn))
The :where clauses are data patterns [entity attribute value]. A symbol
starting with ? is a variable, _ matches anything, and a trailing
position can be left out. Clauses run in the order written, in nitomic as in
Datomic, so put the most selective clause first.
Find specs
The shape of the result is set by the find spec:
| find spec | returns | example |
|---|---|---|
:find ?a ?b | a set of tuples | #{[1 "x"] [2 "y"]} |
:find [?a ...] | a collection of values | ["x" "y"] |
:find [?a ?b] | a single tuple | [1 "x"] |
:find ?a . | a single value | "x" |
The collection form is handy for lists of names:
(d/q '[:find [?n ...] :where [_ :community/name ?n]] (d/db conn))
community names, coll find: ("15th Ave Community" "Admiral Neighborhood Association" ...)
Note the result has no duplicates: several communities share a name (there are three "Magnolia Voice" entries, one per medium), but the collection find returns each name once. The entity-based "community names" listing in the previous chapter showed all 150.
Constants in patterns
A pattern can fix any position to a constant:
(d/q '[:find [?c ...]
:where
[?e :community/name "belltown"]
[?e :community/category ?c]]
(d/db conn))
belltown categories: ("events" "news")
An enum can be given by its keyword in the value position of a ref attribute:
(d/q '[:find [?n ...]
:where
[?c :community/name ?n]
[?c :community/type :community.type/twitter]]
(d/db conn))
twitter feeds: ("Columbia Citizens" "Discover SLU" "Fremont Universe" "Magnolia Voice" "Maple Leaf Life" "MyWallingford")
Joins
When a variable appears in several clauses, the clauses join on it. This query walks from community to neighborhood to district to region:
(d/q '[:find [?c_name ...]
:where
[?c :community/name ?c_name]
[?c :community/neighborhood ?n]
[?n :neighborhood/district ?d]
[?d :district/region :region/ne]]
(d/db conn))
NE region: ("Aurora Seattle" "Hawthorne Hills Community Website" "KOMO Communities - U-District" "KOMO Communities - View Ridge" "Laurelhurst Community Club" "Magnuson Community Garden" "Magnuson Environmental Stewardship Alliance" "Maple Leaf Community Council" "Maple Leaf Life")
To get an enum back as a keyword, join through :db/ident:
(d/q '[:find ?c_name ?r_name
:where
[?c :community/name ?c_name]
[?c :community/neighborhood ?n]
[?n :neighborhood/district ?d]
[?d :district/region ?r]
[?r :db/ident ?r_name]]
(d/db conn))
names and regions: (["15th Ave Community" :region/e] ["Admiral Neighborhood Association" :region/sw] ...)
Parameters with :in
:in names the inputs. $ is the database; other names bind the extra
arguments to d/q. A query with a parameter can be defined once and reused:
(def query-by-type '[:find [?n ...]
:in $ ?t
:where
[?c :community/name ?n]
[?c :community/type ?t]])
(d/q query-by-type (d/db conn) :community.type/twitter)
(d/q query-by-type (d/db conn) :community.type/facebook-page)
by type: twitter: ("Columbia Citizens" "Discover SLU" "Fremont Universe" "Magnolia Voice" "Maple Leaf Life" "MyWallingford")
by type: facebook: ("Blogging Georgetown" "Columbia Citizens" "Discover SLU" "Eastlake Community Council" "Fauntleroy Community Association" "Fremont Universe" "Magnolia Voice" "Maple Leaf Life" "MyWallingford")
The same works with a pull in the find spec:
(def query-by-type-with-pull '[:find (pull ?c [:community/name])
:in $ ?t
:where
[?c :community/type ?t]])
by type with pull: twitter: ([{:community/name "Columbia Citizens"}] [{:community/name "Discover SLU"}] ...)
Binding forms
An input can be destructured:
| binding | binds | input |
|---|---|---|
?x | a scalar | :community.type/twitter |
[?x ?y] | a tuple | [:a :b] |
[?x ...] | a collection, one match per element | [:a :b :c] |
[[?x ?y]] | a relation, one match per tuple | [[:a 1] [:b 2]] |
A collection binding acts as an "or" over its values:
(d/q '[:find ?n ?t
:in $ [?t ...]
:where
[?c :community/name ?n]
[?c :community/type ?t]]
(d/db conn)
[:community.type/facebook-page :community.type/twitter])
collection input: (["Blogging Georgetown" :community.type/facebook-page] ["Columbia Citizens" :community.type/facebook-page] ["Columbia Citizens" :community.type/twitter] ...)
A relation binding matches whole tuples, here pairs of type and orgtype:
(d/q '[:find ?n ?t ?ot
:in $ [[?t ?ot]]
:where
[?c :community/name ?n]
[?c :community/type ?t]
[?c :community/orgtype ?ot]]
(d/db conn)
[[:community.type/email-list :community.orgtype/community]
[:community.type/website :community.orgtype/commercial]])
relation input: (["15th Ave Community" :community.type/email-list :community.orgtype/community] ... ["Discover SLU" :community.type/website :community.orgtype/commercial] ...)
Predicates and functions
A clause of the form [(f args...)] is a predicate: it keeps the matches
for which it returns true. A clause [(f args...) ?out] is a function call
whose result is bound to ?out. Here, .compareTo computes an ordering
and < filters on it:
(d/q '[:find [?n ...]
:where
[?c :community/name ?n]
[(.compareTo ?n "C") ?res]
[(< ?res 0)]]
(d/db conn))
names before C: ("15th Ave Community" "Admiral Neighborhood Association" ... "Blogging Georgetown" "Broadview Community Council")
nitomic: there is no runtime code loading, so query functions come from a built-in table of
clojure.coreand string functions (including.compareTo,.startsWithand the like). You can also pass a function as a query input, or register one by name withd/register-fn!. See Differences from Datomic.
Fulltext search
fulltext searches an attribute declared with :db/fulltext true. It
takes the database, the attribute and the search string, and binds a
relation of [entity value] (Datomic also offers tx and score):
(d/q '[:find ?n .
:where
[(fulltext $ :community/name "Wallingford") [[?e ?n]]]]
(d/db conn))
fulltext Wallingford: "KOMO Communities - Wallingford"
Fulltext combines with ordinary clauses and inputs:
(d/q '[:find ?name ?cat
:in $ ?type ?search
:where
[?c :community/name ?name]
[?c :community/type ?type]
[(fulltext $ :community/category ?search) [[?c ?cat]]]]
(d/db conn)
:community.type/website
"food")
fulltext food websites: (["Community Harvest of Southwest Seattle" "sustainable food"] ["InBallard" "food"])
nitomic: fulltext tokenizes on letters and digits, lower-cases, drops English stop words, and matches any query term (
term*matches a prefix). On typical text that matches Lucene's default analyzer, but it is not Lucene's full query syntax, and every score is 1.0.
Rules
A rule is a named group of :where clauses. You pass a set of rules to a
query as the input named %, and call a rule like a clause.
A simple rule
(let [rules '[[[twitter ?c]
[?c :community/type :community.type/twitter]]]]
(d/q '[:find [?n ...]
:in $ %
:where
[?c :community/name ?n]
(twitter ?c)]
(d/db conn)
rules))
rule: twitter: ("Columbia Citizens" "Discover SLU" "Fremont Universe" "Magnolia Voice" "Maple Leaf Life" "MyWallingford")
The rule set is a vector of rules. Each rule is a vector whose first element
is the head, [name ?args...], followed by the body clauses.
Rules with arguments
A rule packages up a join so queries don't have to repeat it. The region
rule relates a community to the keyword of its region:
(let [rules '[[[region ?c ?r]
[?c :community/neighborhood ?n]
[?n :neighborhood/district ?d]
[?d :district/region ?re]
[?re :db/ident ?r]]]]
(d/q '[:find [?n ...]
:in $ %
:where
[?c :community/name ?n]
[region ?c :region/ne]]
(d/db conn)
rules))
rule: NE: ("Aurora Seattle" "Hawthorne Hills Community Website" ... "Maple Leaf Life")
rule: SW: ("Admiral Neighborhood Association" "Alki News" ... "Nature Consortium")
A rule call may be written in brackets, [region ?c :region/ne], or in
parentheses, (region ?c :region/ne); both mean the same. An argument can
be a variable or a constant.
Or, by defining a rule more than once
Several rules with the same head are alternatives: an entity matches if any
of them matches. Rules can also call other rules. This set builds
social-media, northern and southern on top of region:
(let [rules '[[[region ?c ?r]
[?c :community/neighborhood ?n]
[?n :neighborhood/district ?d]
[?d :district/region ?re]
[?re :db/ident ?r]]
[[social-media ?c]
[?c :community/type :community.type/twitter]]
[[social-media ?c]
[?c :community/type :community.type/facebook-page]]
[[northern ?c]
(region ?c :region/ne)]
[[northern ?c]
(region ?c :region/n)]
[[northern ?c]
(region ?c :region/nw)]
[[southern ?c]
(region ?c :region/sw)]
[[southern ?c]
(region ?c :region/s)]
[[southern ?c]
(region ?c :region/se)]]]
(d/q '[:find [?n ...]
:in $ %
:where
[?c :community/name ?n]
(southern ?c)
(social-media ?c)]
(d/db conn)
rules))
rule: southern social media: ("Blogging Georgetown" "Columbia Citizens" "Fauntleroy Community Association" "MyWallingford")
nitomic also supports recursive rules, including over cyclic data, as well
as or, or-join, and, not and not-join clauses inside queries and
rules.
Time travel
A Datomic database is an accumulation of facts, and a database value is
immutable. Every transaction is itself an entity, with a :db/txInstant
recording when it happened. That makes it possible to look at the database
as it was at any point, or at only what changed since then.
Finding transaction times
(def tx-instants (reverse (sort (d/q '[:find [?when ...] :where [_ :db/txInstant ?when]]
(d/db conn)))))
(show "transaction instants" (count tx-instants))
(def data-tx-date (first tx-instants))
(def schema-tx-date (second tx-instants))
transaction instants: 3
There are three transactions: the bootstrap transaction that every database starts with, the schema, and the data. Sorted newest first, the first instant is the data transaction and the second is the schema transaction.
The rest of this chapter runs one query against different views of the database:
(def communities-query '[:find [?c ...] :where [?c :community/name]])
as-of: the database at a point in time
d/as-of returns the database as it was at a time, which can be given as a
t, a transaction id or an instant:
(let [db-asof-schema (-> conn d/db (d/as-of schema-tx-date))]
(count (d/q communities-query db-asof-schema)))
(let [db-asof-data (-> conn d/db (d/as-of data-tx-date))]
(count (d/q communities-query db-asof-data)))
as of schema: 0
as of data: 150
Right after the schema transaction there were no communities yet; right after the data transaction there were 150.
since: only what changed after a point
d/since returns a database that contains only the facts added after a
time:
(let [db-since-data (-> conn d/db (d/since schema-tx-date))]
(count (d/q communities-query db-since-data)))
(let [db-since-data (-> conn d/db (d/since data-tx-date))]
(count (d/q communities-query db-since-data)))
since schema: 150
since data: 0
with: what if?
d/with applies a transaction to a database value without committing it. It
returns the same kind of report as transact, and its :db-after is a
database you can query. The connection is untouched:
(def new-data-tx (read-edn "examples/seattle/seattle-data1.edn"))
(let [db-if-new-data (-> conn d/db (d/with new-data-tx) :db-after)]
(count (d/q communities-query db-if-new-data)))
(count (d/q communities-query (d/db conn)))
with new data: 258
current: 150
This is useful for trying a transaction out, validating it, or computing a speculative result.
Committing the new data
Now transact it for real:
@(d/transact conn new-data-tx)
(count (d/q communities-query (d/db conn)))
(let [db-since-data (-> conn d/db (d/since data-tx-date))]
(count (d/q communities-query db-since-data)))
after new data: 258
since first data: 108
The database now has 258 communities, and since shows exactly the 108 that
the new transaction added.
Note that seattle-data1.edn mentions neighborhoods and districts that
already exist, such as "Beacon Hill". Because their names are unique
identities, those tempids upsert onto the existing entities instead of
creating duplicates.
More time functions
nitomic also supports d/history (a database of every assertion and
retraction ever made), d/as-of-t, d/since-t, d/basis-t, d/next-t,
d/is-history, d/filter and d/is-filtered, and the transaction log via
d/log and d/tx-range.
Changing data
The last part of the walkthrough makes smaller changes: it creates a partition, adds and retracts values, retracts a whole entity, and watches transactions through the report queue.
Creating a partition
Partitions group entities, which in Datomic affects how their datoms sort
together. A partition is an entity with a :db/ident, installed through
:db.install/_partition:
(def partition-report
@(d/transact conn [{:db/id (d/tempid :db.part/db)
:db/ident :communities
:db.install/_partition :db.part/db}]))
(show "partition installed" (contains? (set (map :v (:tx-data partition-report)))
:communities))
partition installed: true
d/tempid creates a tempid in a partition, like a #db/id literal does.
:db.install/_partition :db.part/db is a reverse attribute in a map: it
asserts [:db.part/db :db.install/partition <new entity>].
Creating an entity in the partition
(def easton-report
@(d/transact conn [{:db/id (d/tempid :communities)
:community/name "Easton"}]))
(def easton-created (d/q '[:find ?id . :where [?id :community/name "Easton"]] (d/db conn)))
(show "Easton partition is :communities"
(= (d/part easton-created) (d/entid (d/db conn) :communities)))
Easton partition is :communities: true
d/part returns the partition of an entity id, and d/entid turns an ident
into its entity id.
Adding a value
To add to an existing entity, transact a map with its :db/id:
(def belltown-id (d/q '[:find ?id .
:where
[?id :community/name "belltown"]]
(d/db conn)))
@(d/transact conn [{:db/id belltown-id
:community/category "free stuff"}])
(:community/category (d/entity (d/db conn) belltown-id))
belltown categories after add: ("events" "free stuff" "news")
:community/category is cardinality many, so the new value joins the
existing ones. For a cardinality-one attribute, the new value would replace
the old one (Datomic retracts the old value for you).
Retracting a value
The list form [:db/retract e a v] retracts one value:
@(d/transact conn [[:db/retract belltown-id :community/category "free stuff"]])
(:community/category (d/entity (d/db conn) belltown-id))
belltown categories after retract: ("events" "news")
The list form [:db/add e a v] is the counterpart for assertions.
Retracting an entity
:db.fn/retractEntity (also spelled :db/retractEntity) retracts every
attribute of an entity, and every reference to it:
(def easton-id (d/q '[:find ?id .
:where
[?id :community/name "Easton"]]
(d/db conn)))
@(d/transact conn [[:db.fn/retractEntity easton-id]])
(d/q '[:find ?id . :where [?id :community/name "Easton"]] (d/db conn))
Easton after retractEntity: nil
A retraction is a new fact, not a deletion: the history still records that
Easton existed, and d/as-of an earlier time still finds it.
Watching transactions: the tx report queue
d/tx-report-queue returns a queue that receives the report of every
transaction committed on the connection after it was created:
(def queue (d/tx-report-queue conn))
@(d/transact conn [{:db/id (d/tempid :communities)
:community/name "Easton"}])
.poll takes the next report, or returns nil if there is none. The report
has the same keys as a transact result. The walkthrough queries its
:tx-data directly: a collection of datoms can be a query input, bound here
as a relation of [e a v tx added]:
(when-let [report (.poll queue)]
(show-sorted "tx report from queue"
(map (fn [[e aname v added]]
[(if (= 3 (d/part e)) :tx (d/part e)) aname
(if (inst? v) :inst v) added])
(d/q '[:find ?e ?aname ?v ?added
:in $ [[?e ?a ?v _ ?added]]
:where
[?e ?a ?v _ ?added]
[?a :db/ident ?aname]]
(:db-after report)
(:tx-data report)))))
(show "queue empty after poll" (.poll queue))
tx report from queue: ([83 :community/name "Easton" true] [:tx :db/txInstant :inst true])
queue empty after poll: nil
The transaction produced two datoms: the new community's name, on an entity
in partition 83 (the :communities partition), and the transaction's own
:db/txInstant, on the transaction entity in partition 3. After one poll the
queue is empty.
nitomic's queue also supports .take, .peek, .isEmpty and .size, and
d/remove-tx-report-queue detaches it.
That's the walkthrough
The program ends with (System/exit 0). On the JVM this shuts down
Datomic's background threads; natively it simply exits.
From here, the What works page lists the rest of
the API, and test/features.clj in the repository exercises it, including
the error cases.
What works
The table below lists the parts of datomic.api that nitomic implements.
For what is missing or behaves differently, see
Differences from Datomic.
| Area | Supported |
|---|---|
| Connections | create-database connect delete-database rename-database get-database-names db release shutdown sync request-index |
| Transactions | transact transact-async with; map and list forms; nested maps; reverse attributes in maps; cardinality-many values; :db/add :db/retract (with or without a value) :db/retractEntity :db/cas (and their :db.fn/ spellings); datoms as tx data |
| Ids | tempid (#db/id literals via datomic.db/id-literal), string tempids, implicit tempids, resolve-tempid, upsert through :db.unique/identity (including between tempids in one transaction), lookup refs everywhere, idents, entid ident entid-at part t->tx tx->t squuid squuid-time-millis |
| Schema | implicit attribute installation, :db/unique (value and identity), :db/isComponent, :db/index, :db/fulltext, :db/noHistory, enums, partitions (:db.install/partition), schema alteration, attribute |
| Values | string, long, double/float, boolean, keyword, symbol, ref, instant (#inst), uuid (#uuid), uri, bigint/bigdec (as numbers), tuple (as vectors), fn |
| Query | q query, list/map/string queries; find specs rel, [?x ...], [?a ?b], ?x .; pull in :find; :with; :in scalars, tuples, collections, relations, several sources, rules (%); data patterns incl. tx and added positions; predicates and functions with all binding forms; not not-join or or-join and; rules, recursive and over cyclic data; aggregates count count-distinct sum min max avg median distinct (min n ?x) (max n ?x) sample rand; get-else get-some missing? ground tuple untuple fulltext |
| Pull | pull pull-many; *, :db/id, reverse attributes, nested maps, recursion (... and depth limits), :as :limit :default (vector or list syntax) and legacy (limit ...) (default ...), component attributes pulled recursively |
| Entities | entity touch entity-db; lazy lookup, keyword/get access, keys, reverse navigation (:ns/_attr), enums as keywords, cardinality-many as sets |
| Time | as-of since history (by t, tx id or instant), as-of-t since-t basis-t next-t is-history filter is-filtered |
| Indexes | datoms seek-datoms index-range over :eavt :aevt :avet :vaet; datoms support :e :a :v :tx :added and nth |
| Log | log tx-range tx-report-queue (.poll .take .peek .isEmpty .size) remove-tx-report-queue |
| Errors | ex-info with Datomic's :db/error codes (:db.error/unique-conflict, :db.error/cas-failed, :db.error/wrong-type-for-attribute, :db.error/not-an-entity, :db.error/datoms-conflict, …); a failed transact throws when dereferenced |
Durable storage
A datomic:sql URI with a JDBC URL makes a database durable, in SQLite or
in PostgreSQL:
;; one machine: a SQLite file
(def uri "datomic:sql://hello?jdbc:sqlite:/var/lib/app/datomic.db")
;; any number of machines: a PostgreSQL database, from any provider
(def uri (str "datomic:sql://hello?jdbc:postgresql://db.example.com:5432/app"
"?user=app&password=" (System/getenv "PGPASSWORD") "&sslmode=require"))
(d/create-database uri)
(def conn (d/connect uri))
The storage is created on first use and can hold any number of databases.
create-database, delete-database, rename-database and
get-database-names (with datomic:sql://*?<jdbc-url>) act on its catalog.
A PostgreSQL URL is handed to libpq without its jdbc: prefix, so anything
libpq accepts works, TLS options included.
- What is stored. Each transaction is stored as one log row: the datoms
it produced, its tempids and the id counters after it.
connectreplays the log to rebuild the database, so every id and value comes back as the transaction made it. The transaction logic never runs twice. - Several processes. Any number of processes can share a storage.
transacttakes the storage's write lock, applies what others have committed since it last looked, transacts, and stores the result before releasing the lock. The lock isBEGIN IMMEDIATEon SQLite and a transaction-scoped advisory lock on PostgreSQL. So each process acts as its own transactor, one at a time. - Two round trips per write on PostgreSQL. A write sends two multi-statement queries. The first begins, takes the lock, checks for a transactor, and reads what others committed. The second appends, notifies, and commits. Latency to the server matters more than anything else here, so keep the database in the same region as the peers.
- Seeing other writers.
d/db,syncand reading atx-report-queuepick up other writers' transactions, with their:tempids. - Push. On PostgreSQL every commit sends a
NOTIFY, so a process waiting for a transaction wakes as soon as it lands. That covers dereferencing a queued transaction and atx-report-queue's.take, which blocks until the next transaction. SQLite has no way to tell other processes, so there waiting means looking again every couple of milliseconds, and.takereturns nil when the queue is empty. - Transaction functions. A
:db/fnholds a Clojure fn, which can't be stored. Installing one in a stored database fails with:db.error/not-storable. - Releasing.
releasedrops a stored connection from the cache, so the nextconnectrebuilds it from storage. - Transactor. Running a transactor makes it the only process that writes the log (see Running a transactor).
- Requirements. Storage uses clonim's
clonim.sqliteandclonim.postgres. They loadlibsqlite3orlibpqwhen a stored database is first used.
Running a transactor
clonim build script/transactor.clj --source-path src -o nitomic-transactor
./nitomic-transactor 'jdbc:postgresql://host/app?user=app&password=...'
./nitomic-transactor /var/lib/app/datomic.db # SQLite
NITOMIC_STORAGE='jdbc:postgresql://...' ./nitomic-transactor
nitomic.transactor serves every database in a storage. It claims the
storage while it runs:
- The claim. On PostgreSQL the claim is a session advisory lock, which the server drops the moment the transactor's connection goes, even if it is killed. SQLite has nothing like that, so there the transactor records a heartbeat every second instead.
- Queued transactions. While a transactor has the storage,
d/transactdoesn't write the log. It queues the transaction data and returns a future at once. The transactor takes queued transactions in order. For each one, in a single transaction, it runs it, appends it to the log, and records the outcome. Dereferencing the future waits for that outcome, then returns the report or throws the transactor's error. - Waiting. On PostgreSQL the transactor sleeps until a peer's
NOTIFYsays something was queued, and peers sleep until itsNOTIFYsays their transaction is done. - Clock. The transactor stamps
:db/txInstantwith its own clock, as Datomic's does. - Only one at a time. A second transactor refuses to start
(
:db.error/transactor-running). - When it stops. Peers go back to writing the log themselves, so the
databases stay writable. On PostgreSQL that happens at once; on SQLite,
after three missed heartbeats. A transaction still waiting in the queue at
that point is withdrawn and fails with
:db.error/transactor-unavailable. - What can be queued. Transaction data crosses processes as EDN, with
tempids and datoms encoded. A fn can't be sent and fails with
:db.error/not-storable. - Programmatic use.
nitomic.transactor/runserves a storage until its:stop?fn returns true.start,step!andstop!drive it one batch at a time.
Deploying on Modal
On Modal every container is its own machine, and
containers come and go, so use PostgreSQL storage from any provider (Neon,
Supabase, RDS, your own). A Modal Volume can't hold a shared SQLite file:
its changes only reach other containers through commit() and reload(),
which SQLite's locking knows nothing about.
- Build. Build each program with
clonim build, and copy the binaries into an image that haslibpq5andlibpcre3installed. - Credentials. Keep the connection URL in a Modal Secret and read it
with
System/getenv. - Direct connections. Use the provider's direct connection string, not
a transaction-mode pooler (PgBouncer, Neon's
-poolerhost, Supabase's port 6543). A pooler hands each transaction to a different server session, which breaksLISTENand the transactor's session lock. - Peers. Run your app as usual (for example behind
@modal.web_server). Every container is a peer: it connects, replays the log once, and then catches up. - Transactor (optional). Run
nitomic-transactoras a single always-on function (min_containers=1,max_containers=1). If Modal restarts it, peers write the log themselves until it is back.
The repository's deploy/modal/ directory does all of this. It has a Modal
app with an HTTP API over nitomic peers and an optional transactor, deployed
with modal deploy deploy/modal/app.py. With Neon as the
PostgreSQL, see deploy/modal/README.md: use Neon's direct (not pooled)
connection string, and deploy without the transactor
(NITOMIC_TRANSACTOR=0) if the compute should scale to zero.
Differences from Datomic
- Storage is memory, SQLite or PostgreSQL.
datomic:sqlURIs with ajdbc:sqlite:orjdbc:postgresql:URL keep databases durably (see Durable storage). Every other URI protocol (mem,dev,ddb, …) names an in-process database that is gone when the process exits. The transactor is optional: without one, each process takes the storage's write lock to transact. - No runtime code compilation. Database functions can't be Clojure source
strings.
:db/fnholds a Clojure fn, whichd/functionpasses through. Query functions are found in a built-in table ofclojure.coreand string functions (including.compareTo,.startsWith, and the like). You can also pass them as inputs or register them withd/register-fn!. The JVM would resolve any qualified symbol instead. Seetest/native.clj. - Fulltext tokenizes on letters and digits, lower-cases, drops English
stop words, and matches any query term (with
term*prefixes). That matches the default Lucene analyzer on typical text. It is not Lucene's full query syntax, and scores are all 1.0. - Not implemented: composite tuple attributes (
:db/tupleAttrs),:db/ensureand entity specs, excision,d/index-pull, and database stats beyond counts. - Printing. Query results are Clojure sets and vectors rather than Java
collections. An entity prints as
{:db/id n}, and a touched entity prints as its attribute map.
How it is tested
nitomic is tested by comparison with Datomic itself. Each test program uses
only datomic.api, runs unchanged on the JVM against Datomic Pro and
natively against nitomic, and prints its results in a canonical order. The
native output must match the recorded JVM output line for line.
The test programs
| program | what it covers | expected output |
|---|---|---|
examples/seattle/getting_started.clj | Datomic's getting-started walkthrough over the Seattle data (the Getting started part of this book) | recorded on Datomic Pro |
test/features.clj | the rest of the API, including error cases | recorded on Datomic Pro |
test/native.clj | nitomic-only extensions, such as d/register-fn! | written for nitomic |
test/storage.clj | durable storage (SQLite, or PostgreSQL via NITOMIC_TEST_STORAGE): replay, two connections sharing a storage, the catalog | written for nitomic |
test/transactor.clj | the transactor: queued transactions, errors, tempids, stopping and withdrawal | written for nitomic |
Running the tests
CLONIM=path/to/clonim/bin/clonim test/run.sh
# or, with clonim on PATH
nimble test
test/run.sh runs each program with clonim run ... --source-path src,
diffs its output against test/expected/<name>.out, and prints ok or
FAIL with the first lines of the diff:
ok getting_started
ok features
ok native
ok storage
ok transactor
The storage and transactor tests use a SQLite file unless
NITOMIC_TEST_STORAGE names another storage. Given a PostgreSQL URL, they
must print exactly the same output.
CI (.github/workflows/test.yml) builds clonim with Nim 2.2.12 and runs the
same script on every push and pull request twice: once as is, and once
against a PostgreSQL service container.
Re-recording the expected output
The expected outputs of the first two programs were recorded by
script/reference.sh on the JVM against Datomic Pro 1.0.7705. To re-record
them from a distribution:
curl -O https://datomic-pro-downloads.s3.amazonaws.com/1.0.7705/datomic-pro-1.0.7705.zip
unzip datomic-pro-1.0.7705.zip
script/reference.sh datomic-pro-1.0.7705
The script runs each program with clojure.main on the peer jar, and drops
the JVM's log lines and reflection warnings so that only program output is
kept.
Writing a comparable program
To make a new program comparable across the two platforms, follow the walkthrough's conventions:
- print results through a canonical printer (see the
showhelpers in the Overview), so that set order doesn't matter; - print instants as a placeholder and leave tempids out, since they differ between runs and platforms;
- end with
(System/exit 0)so the JVM run terminates.
Source layout
| file | what it does |
|---|---|
src/datomic/api.clj | the public API |
src/datomic/db.clj | id-literal, the #db/id reader function |
src/nitomic/db.clj | database values: covering indexes as nested persistent maps, the schema cache, as-of/since/history views |
src/nitomic/tx.clj | transactions: expansion, tempid resolution and upsert, id allocation, index updates, schema installation |
src/nitomic/query.clj | Datalog, rules and aggregates |
src/nitomic/pull.clj | the pull API |
src/nitomic/entity.clj | lazy entities (a deftype over ILookup/Seqable) |
src/nitomic/storage.clj | durable storage in SQLite or PostgreSQL: the catalog, transaction log and queue, replay, locks and notifications |
src/nitomic/transactor.clj | the transactor: serves a storage's queue and writes its log |
src/nitomic/types.clj | datoms and tempids |
src/nitomic/bootstrap.clj | Datomic's bootstrap datoms |
A database is an immutable map, so every database value stays valid. Its
covering indexes ({e {a {v tx}}}, {a {e {v tx}}}, {a {v {e tx}}} and
{v {a {e tx}}} for refs) make every bound prefix a hash lookup. The sorted
orders that d/datoms promises are produced on demand. As in Datomic, query
clauses run in the order written.