Introduction

nitomic is Datomic's peer API (datomic.api) ported to clonim, a Clojure compiler that targets Nim. Datomic programs written against datomic.api compile to native binaries, with no JVM involved.

(require '[datomic.api :as d])

(d/create-database "datomic:mem://hello")
(def conn (d/connect "datomic:mem://hello"))

@(d/transact conn [{:db/ident :person/name
                    :db/valueType :db.type/string
                    :db/cardinality :db.cardinality/one}])
@(d/transact conn [{:person/name "Ada"} {:person/name "Grace"}])

(d/q '[:find [?n ...] :where [_ :person/name ?n]] (d/db conn))
;=> ["Ada" "Grace"]

A clean-room port, checked against the real thing

Datomic Pro 1.0.7705 ships its peer library as AOT-compiled JVM classes, with no Clojure source. nitomic is therefore a clean-room reimplementation of the peer API in Clojure that clonim can compile.

It is checked against Datomic itself. The same test programs run on the JVM against Datomic Pro and natively against nitomic, and their outputs must match line for line. The main one of those programs is Datomic's own getting-started walkthrough, which is what the Getting started part of this book explains step by step.

The bootstrap database is also taken from Datomic. The system partitions, value types and attributes, with Datomic's own entity ids, docs and transaction ids, were dumped from a fresh datomic:mem database. New entity ids are allocated the way Datomic allocates them:

  • transaction ids are t in partition 3;
  • other ids share the t counter in their partition;
  • attributes get their own counter in the db partition.

So the ids a program sees are the ids Datomic would give it.

How to read this book

  • Installation and Your first program get a program compiled and running.
  • Getting started walks through the Seattle example in examples/seattle/, explaining each step and showing the output it produces.
  • The Reference part lists what is supported, where nitomic differs from Datomic, and how the port is tested.

Installation

nitomic is Clojure source, not a compiled library. You need clonim to compile programs, and you point clonim at nitomic's src/ directory with --source-path.

Requirements

  • Nim 2.2.12 or later, with Nimble.
  • clonim, at a revision from main that includes codegod100/clonim#1. That change added the Clojure features nitomic relies on: reify/deftype/defprotocol, #inst/#uuid and *data-readers*, compare and comparator sorts, for/doseq modifiers, ex-data, and a set of collection functions.
  • On Linux, the PCRE runtime library (libpcre3 on Debian and Ubuntu).

If you don't have Nim yet, choosenim is the quickest route:

curl https://nim-lang.org/choosenim/init.sh -sSf | sh
export PATH="$HOME/.nimble/bin:$PATH"

With Nimble

nitomic is a Nimble package, and installing it also installs clonim:

nimble install https://github.com/codegod100/nitomic

Nimble installs src/ as is (the package keeps only .clj files), so the package directory is itself the source path:

clonim run hello.clj --source-path "$(nimble path nitomic)"

From a checkout

git clone https://github.com/codegod100/nitomic
cd nitomic
nimble install          # installs the working tree (and clonim)

A checkout also gives you two Nimble tasks:

taskwhat it does
nimble exampleruns the Seattle walkthrough, examples/seattle/getting_started.clj
nimble testruns every test program and diffs its output against test/expected/

Both use the clonim on your PATH. Set CLONIM to use a different one:

CLONIM=path/to/clonim/bin/clonim nimble example

Building clonim by hand

This is what CI does, and it is handy when you want a specific clonim revision:

git clone https://github.com/codegod100/clonim
cd clonim
nim c --hints:off --warnings:off -o:bin/clonim src/clonim.nim

Then use path/to/clonim/bin/clonim wherever this book says clonim.

Your first program

Save this as hello.clj:

(require '[datomic.api :as d])

(d/create-database "datomic:mem://hello")
(def conn (d/connect "datomic:mem://hello"))

;; Install one attribute.
@(d/transact conn [{:db/ident :person/name
                    :db/valueType :db.type/string
                    :db/cardinality :db.cardinality/one}])

;; Add two people.
@(d/transact conn [{:person/name "Ada"} {:person/name "Grace"}])

(prn (sort (d/q '[:find [?n ...] :where [_ :person/name ?n]] (d/db conn))))

Run it

clonim run hello.clj --source-path path/to/nitomic/src

clonim run compiles the program through Nim and runs it. It prints:

("Ada" "Grace")

The query returns its results in no particular order, so the program sorts them before printing.

Build a binary

clonim build hello.clj --source-path path/to/nitomic/src

clonim build produces a standalone native executable instead of running the program. Release builds (-d:release) are what you want for speed: built this way, the full Seattle walkthrough runs in about 1.3 seconds.

What just happened

The program uses nothing but datomic.api, so the same file also runs on the JVM against Datomic Pro. A few things to notice:

  • datomic:mem://hello names an in-memory database, gone when the program exits. datomic:sql://hello?jdbc:sqlite:hello.db would keep it in a SQLite file instead (see Durable storage).
  • d/transact returns a future. Dereferencing it with @ waits for the result, a map with :db-before, :db-after, :tx-data and :tempids. A failed transaction throws when you dereference it.
  • Maps without :db/id get implicit tempids. Each map in the second transaction becomes a new entity.
  • [?n ...] is a collection find spec: the query returns a vector of names instead of a set of tuples.

The Getting started walkthrough covers all of this, and much more, on a realistic data set.

Overview

Datomic Pro ships a getting-started tutorial built around a small data set about Seattle's neighborhood communities: blogs, mailing lists, Twitter feeds, chambers of commerce and so on. The tutorial lives in the Datomic distribution as samples/seattle/getting-started.clj, a file of forms meant to be evaluated one at a time in a REPL.

nitomic carries that tutorial as a single program:

examples/seattle/
├── getting_started.clj   # the walkthrough, as one program
├── seattle-schema.edn    # the schema: 10 attributes and their enums
├── seattle-data0.edn     # the initial data: 150 communities
└── seattle-data1.edn     # more data, added later: 108 more communities

The data files come from the Datomic Pro distribution, which is licensed under the Apache License 2.0.

Running it

From the root of a nitomic checkout:

nimble example
# or, equivalently
clonim run examples/seattle/getting_started.clj --source-path src

The program reads the .edn files by relative path, so run it from the repository root.

Why it prints the way it does

The program runs unchanged on the JVM against Datomic Pro and natively against nitomic, and the two outputs are compared line by line (see How it is tested). Query results are sets, whose iteration order differs between the two implementations, and instants differ from run to run. So instead of printing results directly, the program prints them through two small helpers that produce a canonical form:

(defn- canon-str [x]
  (cond
    (map? x) (str "{" (str/join ", " (sort (map (fn [[k v]] (str (canon-str k) " " (canon-str v))) x))) "}")
    (set? x) (str "#{" (str/join " " (sort (map canon-str x))) "}")
    (vector? x) (str "[" (str/join " " (map canon-str x)) "]")
    (seq? x) (str "(" (str/join " " (map canon-str x)) ")")
    (inst? x) "#inst"
    :else (pr-str x)))

(defn show
  "Print a labelled result; an unordered collection of results is sorted."
  [label x]
  (println (str label ":") (canon-str x)))

(defn show-sorted [label xs]
  (println (str label ":") (str "(" (str/join " " (sort (map canon-str xs))) ")")))
  • show prints a label and a value. Map keys and set members are sorted, and every instant prints as #inst.
  • show-sorted prints a collection of results as a sorted list.

Every output line quoted in the following chapters is taken from test/expected/getting_started.out, which was recorded on the JVM against Datomic Pro 1.0.7705. nitomic must reproduce it exactly.

The chapters

chaptercovers
The Seattle data modelthe schema: communities, neighborhoods, districts, enums
Creating a database and loading datacreate-database, connect, #db/id tempids, transact
Entities and pullentity, navigation and reverse navigation, pull in queries
Queriesfind specs, joins, :in parameters, predicates, fulltext
Rulesnamed, reusable, composable query clauses
Time travelas-of, since, with
Changing datapartitions, adding and retracting values, retractEntity, the tx report queue

The Seattle data model

The schema in examples/seattle/seattle-schema.edn describes three kinds of entity, linked by references:

community ──:community/neighborhood──▶ neighborhood ──:neighborhood/district──▶ district
    │                                                                              │
    ├─ :community/type     ─▶ enum (:community.type/...)                          └─ :district/region ─▶ enum (:region/...)
    └─ :community/orgtype  ─▶ enum (:community.orgtype/...)

Attributes

attributetypecardinalitynotes
:community/namestringonefulltext
:community/urlstringone
:community/neighborhoodrefonea neighborhood
:community/categorystringmanyfulltext
:community/orgtyperefonean orgtype enum
:community/typerefmanytype enums
:neighborhood/namestringone:db.unique/identity
:neighborhood/districtrefonea district
:district/namestringone:db.unique/identity
:district/regionrefonea region enum

Each attribute is installed with a plain map:

{:db/ident :community/name
 :db/valueType :db.type/string
 :db/cardinality :db.cardinality/one
 :db/fulltext true
 :db/doc "A community's name"}

Datomic (and nitomic) install an attribute implicitly when a transaction asserts :db/ident, :db/valueType and :db/cardinality for a new entity; no :db.install/_attribute is needed.

Unique identities

:neighborhood/name and :district/name are :db.unique/identity. When a transaction asserts one of those values for a tempid, and an entity already has that value, the tempid resolves to the existing entity instead of creating a new one. This is called upsert. It is what lets the second batch of data (seattle-data1.edn) mention "Beacon Hill" again without creating a second Beacon Hill.

Fulltext

:community/name and :community/category are :db/fulltext, which makes them searchable with the fulltext query function (see Queries).

Enums

Enumerated values are entities with nothing but a :db/ident. Refs to them can be written as the keyword:

[:db/add #db/id[:db.part/user] :db/ident :community.type/twitter]
enumvalues
:community.orgtype/...community commercial nonprofit personal
:community.type/...email-list twitter facebook-page blog website wiki myspace ning
:region/...n ne e se s sw w nw

In an entity, a ref to an enum reads back as its keyword. In a query, you join through :db/ident to get the keyword, or pass the keyword as an input.

The data

The data files are vectors of maps. Each map carries a #db/id tempid, and references between new entities use the same tempid:

{:district/region :region/e, :db/id #db/id[:db.part/user -1000001], :district/name "East"}
{:db/id #db/id[:db.part/user -1000002], :neighborhood/name "Capitol Hill",
 :neighborhood/district #db/id[:db.part/user -1000001]}
{:community/category ["15th avenue residents"],
 :community/orgtype :community.orgtype/community,
 :community/type :community.type/email-list,
 :db/id #db/id[:db.part/user -1000003],
 :community/name "15th Ave Community",
 :community/url "http://groups.yahoo.com/group/15thAve_Community/",
 :community/neighborhood #db/id[:db.part/user -1000002]}
  • seattle-data0.edn holds 150 communities with their neighborhoods and districts.
  • seattle-data1.edn holds 108 more, added in Time travel.

Creating a database and loading data

Requiring the API

(require '[datomic.api :as d]
         '[datomic.db]
         '[clojure.string :as str])

datomic.api is the whole public API. datomic.db provides id-literal, the reader function behind #db/id[...] tempid literals, which the data files use.

Creating and connecting

(def uri "datomic:mem://seattle")
(show "create-database" (d/create-database uri))
(def conn (d/connect uri))
create-database: true

create-database returns true when it creates a database, and false if one with that name already exists. connect returns a connection, which is what you transact against and take database values from.

nitomic: datomic:mem and every other protocol except sql name an in-memory database in the same process. To keep a database on disk, use a SQLite URI such as datomic:sql://seattle?jdbc:sqlite:seattle.db, or a jdbc:postgresql: one to share it between machines (see Durable storage).

Reading EDN with tempid literals

The schema and data files contain #db/id[:db.part/user] and #db/id[:db.part/user -1000001] forms. To read them, bind *data-readers* so the db/id tag goes to datomic.db/id-literal:

(defn read-edn [path]
  (binding [*data-readers* {'db/id datomic.db/id-literal}]
    (read-string (slurp path))))

Each literal becomes a tempid in the named partition. Two literals with the same negative number are the same tempid, which is how the data files link new entities to each other. A literal without a number is a fresh tempid.

Transacting the schema

(def schema-tx (read-edn "examples/seattle/seattle-schema.edn"))
(show "first schema statement" (first schema-tx))
(def schema-report @(d/transact conn schema-tx))
(show "schema tx-data count" (count (:tx-data schema-report)))
first schema statement: {:db/cardinality :db.cardinality/one, :db/doc "A community's name", :db/fulltext true, :db/ident :community/name, :db/valueType :db.type/string}
schema tx-data count: 75

d/transact returns a future. Dereferencing it (@) waits for the transaction and returns a report map:

keyvalue
:db-beforethe database value before the transaction
:db-afterthe database value after it
:tx-datathe datoms the transaction asserted or retracted
:tempidsa map from tempids to the entity ids they resolved to

The 75 datoms are the 44 attribute definition datoms, one :db.install/attribute datom per attribute (10), the 20 enum idents, and the transaction's own :db/txInstant. nitomic allocates the same entity ids as Datomic, so the same 75 datoms come out.

If the transaction fails, dereferencing throws an ex-info whose data carries a Datomic :db/error code, such as :db.error/unique-conflict.

Transacting the data

(def data-tx (read-edn "examples/seattle/seattle-data0.edn"))
(show "first data statement" (dissoc (first data-tx) :db/id))
(show "second data statement" (dissoc (second data-tx) :db/id :neighborhood/district))
(def data-report @(d/transact conn data-tx))
(show "data tx-data count" (count (:tx-data data-report)))
first data statement: {:district/name "East", :district/region :region/e}
second data statement: {:neighborhood/name "Capitol Hill"}
data tx-data count: 1237

The :db/id keys are dropped before printing because a tempid prints differently on the two platforms.

Some things this one transaction exercises:

  • Map form. Each map asserts all its attributes for one entity.
  • Tempid references. :neighborhood/district #db/id[:db.part/user -1000001] points at the district created earlier in the same transaction.
  • Enum keywords as ref values. :district/region :region/e resolves the keyword to the enum entity.
  • Cardinality many. :community/category ["events" "news"] asserts one datom per value.
  • Upsert. A neighborhood that appears twice under two tempids resolves to one entity, because :neighborhood/name is a unique identity.

With the schema and data in place, the database holds 150 communities.

Entities and pull

There are two ways to get at the attributes of an entity: the entity API, which gives a lazy, navigable map-like object, and pull, which gives plain data in the shape you ask for.

Finding the communities

(def results (d/q '[:find ?c :where [?c :community/name]] (d/db conn)))
(show "communities" (count results))
communities: 150

(d/db conn) returns the current database value, an immutable snapshot. The query finds every entity ?c that has a :community/name; the value position of the pattern is left out, so it matches any value. The result is a set of one-element tuples.

Entities

(def id (ffirst (sort-by first results)))
(def entity (-> conn d/db (d/entity id)))
(show "entity keys" (set (keys entity)))
(show "entity name" (:community/name entity))
entity keys: #{:community/category :community/name :community/neighborhood :community/orgtype :community/type :community/url}
entity name: "15th Ave Community"

d/entity returns an entity for an id. Attributes are fetched lazily when you look them up with a keyword or get. keys lists the attributes the entity has.

The walkthrough sorts the results and takes the smallest id, rather than using ffirst directly, so that both platforms pick the same community.

A ref attribute returns another entity, so you can walk the graph:

(let [db (d/db conn)]
  (show-sorted "names and neighborhoods"
               (map #(let [entity (d/entity db (first %))]
                       [(:community/name entity)
                        (-> entity :community/neighborhood :neighborhood/name)])
                    results)))
names and neighborhoods: (["15th Ave Community" "Capitol Hill"] ["Admiral Neighborhood Association" "Admiral (West Seattle)"] ...)

Reverse navigation

Prefixing the attribute name with _ follows a reference backwards. From a neighborhood, :community/_neighborhood returns every community that points at it, as a set of entities:

(def community (d/entity (d/db conn) (ffirst (sort-by first results))))
(def neighborhood (:community/neighborhood community))
(def communities (:community/_neighborhood neighborhood))
(show-sorted "communities in the same neighborhood" (map :community/name communities))
communities in the same neighborhood: ("15th Ave Community" "CHS Capitol Hill Seattle Blog" "Capitol Hill Community Council" "Capitol Hill Housing" "Capitol Hill Triangle" "KOMO Communities - Captol Hill")

Other entity behaviour worth knowing:

  • a cardinality-many attribute returns a set;
  • a ref to an enum returns the enum's keyword;
  • d/touch loads every attribute, and a touched entity prints as its attribute map;
  • in nitomic, an untouched entity prints as {:db/id n}.

Pull in a query

Put a pull expression in :find to get maps instead of ids:

(def pull-results (d/q '[:find (pull ?c [*]) :where [?c :community/name]] (d/db conn)))
(show "pull results" (count pull-results))
pull results: 150

The pattern [*] pulls every attribute. Refs come back as nested maps holding :db/id, and cardinality-many attributes come back as vectors. Here is the belltown community with its refs reduced to their keys:

a pulled community: {:community/category ["events" "news"], :community/name "belltown", :community/neighborhood (:db/id), :community/orgtype (:db/id), :community/type ((:db/id)), :community/url "http://www.belltownpeople.com/"}

A pull pattern can name just the attributes you want, and a :find can mix pulls with plain variables:

(d/q '[:find ?n (pull ?c [:community/url])
       :where [?c :community/name ?n]]
     (d/db conn))
names with urls: (["15th Ave Community" {:community/url "http://groups.yahoo.com/group/15thAve_Community/"}] ...)

Outside queries, d/pull and d/pull-many take a database, a pattern and an entity id (or ids). Patterns support nested maps for refs, reverse attributes, recursion, and the :as, :limit and :default options.

Queries

d/q takes a query and its inputs, and the first input is usually a database. A query is data: a vector (or a map, or a string) with :find, optional :in and :with, and :where.

(d/q '[:find ?c :where [?c :community/name]] (d/db conn))

The :where clauses are data patterns [entity attribute value]. A symbol starting with ? is a variable, _ matches anything, and a trailing position can be left out. Clauses run in the order written, in nitomic as in Datomic, so put the most selective clause first.

Find specs

The shape of the result is set by the find spec:

find specreturnsexample
:find ?a ?ba set of tuples#{[1 "x"] [2 "y"]}
:find [?a ...]a collection of values["x" "y"]
:find [?a ?b]a single tuple[1 "x"]
:find ?a .a single value"x"

The collection form is handy for lists of names:

(d/q '[:find [?n ...] :where [_ :community/name ?n]] (d/db conn))
community names, coll find: ("15th Ave Community" "Admiral Neighborhood Association" ...)

Note the result has no duplicates: several communities share a name (there are three "Magnolia Voice" entries, one per medium), but the collection find returns each name once. The entity-based "community names" listing in the previous chapter showed all 150.

Constants in patterns

A pattern can fix any position to a constant:

(d/q '[:find [?c ...]
       :where
       [?e :community/name "belltown"]
       [?e :community/category ?c]]
     (d/db conn))
belltown categories: ("events" "news")

An enum can be given by its keyword in the value position of a ref attribute:

(d/q '[:find [?n ...]
       :where
       [?c :community/name ?n]
       [?c :community/type :community.type/twitter]]
     (d/db conn))
twitter feeds: ("Columbia Citizens" "Discover SLU" "Fremont Universe" "Magnolia Voice" "Maple Leaf Life" "MyWallingford")

Joins

When a variable appears in several clauses, the clauses join on it. This query walks from community to neighborhood to district to region:

(d/q '[:find [?c_name ...]
       :where
       [?c :community/name ?c_name]
       [?c :community/neighborhood ?n]
       [?n :neighborhood/district ?d]
       [?d :district/region :region/ne]]
     (d/db conn))
NE region: ("Aurora Seattle" "Hawthorne Hills Community Website" "KOMO Communities - U-District" "KOMO Communities - View Ridge" "Laurelhurst Community Club" "Magnuson Community Garden" "Magnuson Environmental Stewardship Alliance" "Maple Leaf Community Council" "Maple Leaf Life")

To get an enum back as a keyword, join through :db/ident:

(d/q '[:find ?c_name ?r_name
       :where
       [?c :community/name ?c_name]
       [?c :community/neighborhood ?n]
       [?n :neighborhood/district ?d]
       [?d :district/region ?r]
       [?r :db/ident ?r_name]]
     (d/db conn))
names and regions: (["15th Ave Community" :region/e] ["Admiral Neighborhood Association" :region/sw] ...)

Parameters with :in

:in names the inputs. $ is the database; other names bind the extra arguments to d/q. A query with a parameter can be defined once and reused:

(def query-by-type '[:find [?n ...]
                     :in $ ?t
                     :where
                     [?c :community/name ?n]
                     [?c :community/type ?t]])

(d/q query-by-type (d/db conn) :community.type/twitter)
(d/q query-by-type (d/db conn) :community.type/facebook-page)
by type: twitter: ("Columbia Citizens" "Discover SLU" "Fremont Universe" "Magnolia Voice" "Maple Leaf Life" "MyWallingford")
by type: facebook: ("Blogging Georgetown" "Columbia Citizens" "Discover SLU" "Eastlake Community Council" "Fauntleroy Community Association" "Fremont Universe" "Magnolia Voice" "Maple Leaf Life" "MyWallingford")

The same works with a pull in the find spec:

(def query-by-type-with-pull '[:find (pull ?c [:community/name])
                               :in $ ?t
                               :where
                               [?c :community/type ?t]])
by type with pull: twitter: ([{:community/name "Columbia Citizens"}] [{:community/name "Discover SLU"}] ...)

Binding forms

An input can be destructured:

bindingbindsinput
?xa scalar:community.type/twitter
[?x ?y]a tuple[:a :b]
[?x ...]a collection, one match per element[:a :b :c]
[[?x ?y]]a relation, one match per tuple[[:a 1] [:b 2]]

A collection binding acts as an "or" over its values:

(d/q '[:find ?n ?t
       :in $ [?t ...]
       :where
       [?c :community/name ?n]
       [?c :community/type ?t]]
     (d/db conn)
     [:community.type/facebook-page :community.type/twitter])
collection input: (["Blogging Georgetown" :community.type/facebook-page] ["Columbia Citizens" :community.type/facebook-page] ["Columbia Citizens" :community.type/twitter] ...)

A relation binding matches whole tuples, here pairs of type and orgtype:

(d/q '[:find ?n ?t ?ot
       :in $ [[?t ?ot]]
       :where
       [?c :community/name ?n]
       [?c :community/type ?t]
       [?c :community/orgtype ?ot]]
     (d/db conn)
     [[:community.type/email-list :community.orgtype/community]
      [:community.type/website :community.orgtype/commercial]])
relation input: (["15th Ave Community" :community.type/email-list :community.orgtype/community] ... ["Discover SLU" :community.type/website :community.orgtype/commercial] ...)

Predicates and functions

A clause of the form [(f args...)] is a predicate: it keeps the matches for which it returns true. A clause [(f args...) ?out] is a function call whose result is bound to ?out. Here, .compareTo computes an ordering and < filters on it:

(d/q '[:find [?n ...]
       :where
       [?c :community/name ?n]
       [(.compareTo ?n "C") ?res]
       [(< ?res 0)]]
     (d/db conn))
names before C: ("15th Ave Community" "Admiral Neighborhood Association" ... "Blogging Georgetown" "Broadview Community Council")

nitomic: there is no runtime code loading, so query functions come from a built-in table of clojure.core and string functions (including .compareTo, .startsWith and the like). You can also pass a function as a query input, or register one by name with d/register-fn!. See Differences from Datomic.

fulltext searches an attribute declared with :db/fulltext true. It takes the database, the attribute and the search string, and binds a relation of [entity value] (Datomic also offers tx and score):

(d/q '[:find ?n .
       :where
       [(fulltext $ :community/name "Wallingford") [[?e ?n]]]]
     (d/db conn))
fulltext Wallingford: "KOMO Communities - Wallingford"

Fulltext combines with ordinary clauses and inputs:

(d/q '[:find ?name ?cat
       :in $ ?type ?search
       :where
       [?c :community/name ?name]
       [?c :community/type ?type]
       [(fulltext $ :community/category ?search) [[?c ?cat]]]]
     (d/db conn)
     :community.type/website
     "food")
fulltext food websites: (["Community Harvest of Southwest Seattle" "sustainable food"] ["InBallard" "food"])

nitomic: fulltext tokenizes on letters and digits, lower-cases, drops English stop words, and matches any query term (term* matches a prefix). On typical text that matches Lucene's default analyzer, but it is not Lucene's full query syntax, and every score is 1.0.

Rules

A rule is a named group of :where clauses. You pass a set of rules to a query as the input named %, and call a rule like a clause.

A simple rule

(let [rules '[[[twitter ?c]
               [?c :community/type :community.type/twitter]]]]
  (d/q '[:find [?n ...]
         :in $ %
         :where
         [?c :community/name ?n]
         (twitter ?c)]
       (d/db conn)
       rules))
rule: twitter: ("Columbia Citizens" "Discover SLU" "Fremont Universe" "Magnolia Voice" "Maple Leaf Life" "MyWallingford")

The rule set is a vector of rules. Each rule is a vector whose first element is the head, [name ?args...], followed by the body clauses.

Rules with arguments

A rule packages up a join so queries don't have to repeat it. The region rule relates a community to the keyword of its region:

(let [rules '[[[region ?c ?r]
               [?c :community/neighborhood ?n]
               [?n :neighborhood/district ?d]
               [?d :district/region ?re]
               [?re :db/ident ?r]]]]
  (d/q '[:find [?n ...]
         :in $ %
         :where
         [?c :community/name ?n]
         [region ?c :region/ne]]
       (d/db conn)
       rules))
rule: NE: ("Aurora Seattle" "Hawthorne Hills Community Website" ... "Maple Leaf Life")
rule: SW: ("Admiral Neighborhood Association" "Alki News" ... "Nature Consortium")

A rule call may be written in brackets, [region ?c :region/ne], or in parentheses, (region ?c :region/ne); both mean the same. An argument can be a variable or a constant.

Or, by defining a rule more than once

Several rules with the same head are alternatives: an entity matches if any of them matches. Rules can also call other rules. This set builds social-media, northern and southern on top of region:

(let [rules '[[[region ?c ?r]
               [?c :community/neighborhood ?n]
               [?n :neighborhood/district ?d]
               [?d :district/region ?re]
               [?re :db/ident ?r]]
              [[social-media ?c]
               [?c :community/type :community.type/twitter]]
              [[social-media ?c]
               [?c :community/type :community.type/facebook-page]]
              [[northern ?c]
               (region ?c :region/ne)]
              [[northern ?c]
               (region ?c :region/n)]
              [[northern ?c]
               (region ?c :region/nw)]
              [[southern ?c]
               (region ?c :region/sw)]
              [[southern ?c]
               (region ?c :region/s)]
              [[southern ?c]
               (region ?c :region/se)]]]
  (d/q '[:find [?n ...]
         :in $ %
         :where
         [?c :community/name ?n]
         (southern ?c)
         (social-media ?c)]
       (d/db conn)
       rules))
rule: southern social media: ("Blogging Georgetown" "Columbia Citizens" "Fauntleroy Community Association" "MyWallingford")

nitomic also supports recursive rules, including over cyclic data, as well as or, or-join, and, not and not-join clauses inside queries and rules.

Time travel

A Datomic database is an accumulation of facts, and a database value is immutable. Every transaction is itself an entity, with a :db/txInstant recording when it happened. That makes it possible to look at the database as it was at any point, or at only what changed since then.

Finding transaction times

(def tx-instants (reverse (sort (d/q '[:find [?when ...] :where [_ :db/txInstant ?when]]
                                     (d/db conn)))))
(show "transaction instants" (count tx-instants))

(def data-tx-date (first tx-instants))
(def schema-tx-date (second tx-instants))
transaction instants: 3

There are three transactions: the bootstrap transaction that every database starts with, the schema, and the data. Sorted newest first, the first instant is the data transaction and the second is the schema transaction.

The rest of this chapter runs one query against different views of the database:

(def communities-query '[:find [?c ...] :where [?c :community/name]])

as-of: the database at a point in time

d/as-of returns the database as it was at a time, which can be given as a t, a transaction id or an instant:

(let [db-asof-schema (-> conn d/db (d/as-of schema-tx-date))]
  (count (d/q communities-query db-asof-schema)))

(let [db-asof-data (-> conn d/db (d/as-of data-tx-date))]
  (count (d/q communities-query db-asof-data)))
as of schema: 0
as of data: 150

Right after the schema transaction there were no communities yet; right after the data transaction there were 150.

since: only what changed after a point

d/since returns a database that contains only the facts added after a time:

(let [db-since-data (-> conn d/db (d/since schema-tx-date))]
  (count (d/q communities-query db-since-data)))

(let [db-since-data (-> conn d/db (d/since data-tx-date))]
  (count (d/q communities-query db-since-data)))
since schema: 150
since data: 0

with: what if?

d/with applies a transaction to a database value without committing it. It returns the same kind of report as transact, and its :db-after is a database you can query. The connection is untouched:

(def new-data-tx (read-edn "examples/seattle/seattle-data1.edn"))

(let [db-if-new-data (-> conn d/db (d/with new-data-tx) :db-after)]
  (count (d/q communities-query db-if-new-data)))

(count (d/q communities-query (d/db conn)))
with new data: 258
current: 150

This is useful for trying a transaction out, validating it, or computing a speculative result.

Committing the new data

Now transact it for real:

@(d/transact conn new-data-tx)
(count (d/q communities-query (d/db conn)))

(let [db-since-data (-> conn d/db (d/since data-tx-date))]
  (count (d/q communities-query db-since-data)))
after new data: 258
since first data: 108

The database now has 258 communities, and since shows exactly the 108 that the new transaction added.

Note that seattle-data1.edn mentions neighborhoods and districts that already exist, such as "Beacon Hill". Because their names are unique identities, those tempids upsert onto the existing entities instead of creating duplicates.

More time functions

nitomic also supports d/history (a database of every assertion and retraction ever made), d/as-of-t, d/since-t, d/basis-t, d/next-t, d/is-history, d/filter and d/is-filtered, and the transaction log via d/log and d/tx-range.

Changing data

The last part of the walkthrough makes smaller changes: it creates a partition, adds and retracts values, retracts a whole entity, and watches transactions through the report queue.

Creating a partition

Partitions group entities, which in Datomic affects how their datoms sort together. A partition is an entity with a :db/ident, installed through :db.install/_partition:

(def partition-report
  @(d/transact conn [{:db/id (d/tempid :db.part/db)
                      :db/ident :communities
                      :db.install/_partition :db.part/db}]))
(show "partition installed" (contains? (set (map :v (:tx-data partition-report)))
                                       :communities))
partition installed: true

d/tempid creates a tempid in a partition, like a #db/id literal does. :db.install/_partition :db.part/db is a reverse attribute in a map: it asserts [:db.part/db :db.install/partition <new entity>].

Creating an entity in the partition

(def easton-report
  @(d/transact conn [{:db/id (d/tempid :communities)
                      :community/name "Easton"}]))
(def easton-created (d/q '[:find ?id . :where [?id :community/name "Easton"]] (d/db conn)))
(show "Easton partition is :communities"
      (= (d/part easton-created) (d/entid (d/db conn) :communities)))
Easton partition is :communities: true

d/part returns the partition of an entity id, and d/entid turns an ident into its entity id.

Adding a value

To add to an existing entity, transact a map with its :db/id:

(def belltown-id (d/q '[:find ?id .
                        :where
                        [?id :community/name "belltown"]]
                      (d/db conn)))

@(d/transact conn [{:db/id belltown-id
                    :community/category "free stuff"}])
(:community/category (d/entity (d/db conn) belltown-id))
belltown categories after add: ("events" "free stuff" "news")

:community/category is cardinality many, so the new value joins the existing ones. For a cardinality-one attribute, the new value would replace the old one (Datomic retracts the old value for you).

Retracting a value

The list form [:db/retract e a v] retracts one value:

@(d/transact conn [[:db/retract belltown-id :community/category "free stuff"]])
(:community/category (d/entity (d/db conn) belltown-id))
belltown categories after retract: ("events" "news")

The list form [:db/add e a v] is the counterpart for assertions.

Retracting an entity

:db.fn/retractEntity (also spelled :db/retractEntity) retracts every attribute of an entity, and every reference to it:

(def easton-id (d/q '[:find ?id .
                      :where
                      [?id :community/name "Easton"]]
                    (d/db conn)))

@(d/transact conn [[:db.fn/retractEntity easton-id]])
(d/q '[:find ?id . :where [?id :community/name "Easton"]] (d/db conn))
Easton after retractEntity: nil

A retraction is a new fact, not a deletion: the history still records that Easton existed, and d/as-of an earlier time still finds it.

Watching transactions: the tx report queue

d/tx-report-queue returns a queue that receives the report of every transaction committed on the connection after it was created:

(def queue (d/tx-report-queue conn))

@(d/transact conn [{:db/id (d/tempid :communities)
                    :community/name "Easton"}])

.poll takes the next report, or returns nil if there is none. The report has the same keys as a transact result. The walkthrough queries its :tx-data directly: a collection of datoms can be a query input, bound here as a relation of [e a v tx added]:

(when-let [report (.poll queue)]
  (show-sorted "tx report from queue"
               (map (fn [[e aname v added]]
                      [(if (= 3 (d/part e)) :tx (d/part e)) aname
                       (if (inst? v) :inst v) added])
                    (d/q '[:find ?e ?aname ?v ?added
                           :in $ [[?e ?a ?v _ ?added]]
                           :where
                           [?e ?a ?v _ ?added]
                           [?a :db/ident ?aname]]
                         (:db-after report)
                         (:tx-data report)))))
(show "queue empty after poll" (.poll queue))
tx report from queue: ([83 :community/name "Easton" true] [:tx :db/txInstant :inst true])
queue empty after poll: nil

The transaction produced two datoms: the new community's name, on an entity in partition 83 (the :communities partition), and the transaction's own :db/txInstant, on the transaction entity in partition 3. After one poll the queue is empty.

nitomic's queue also supports .take, .peek, .isEmpty and .size, and d/remove-tx-report-queue detaches it.

That's the walkthrough

The program ends with (System/exit 0). On the JVM this shuts down Datomic's background threads; natively it simply exits.

From here, the What works page lists the rest of the API, and test/features.clj in the repository exercises it, including the error cases.

What works

The table below lists the parts of datomic.api that nitomic implements. For what is missing or behaves differently, see Differences from Datomic.

AreaSupported
Connectionscreate-database connect delete-database rename-database get-database-names db release shutdown sync request-index
Transactionstransact transact-async with; map and list forms; nested maps; reverse attributes in maps; cardinality-many values; :db/add :db/retract (with or without a value) :db/retractEntity :db/cas (and their :db.fn/ spellings); datoms as tx data
Idstempid (#db/id literals via datomic.db/id-literal), string tempids, implicit tempids, resolve-tempid, upsert through :db.unique/identity (including between tempids in one transaction), lookup refs everywhere, idents, entid ident entid-at part t->tx tx->t squuid squuid-time-millis
Schemaimplicit attribute installation, :db/unique (value and identity), :db/isComponent, :db/index, :db/fulltext, :db/noHistory, enums, partitions (:db.install/partition), schema alteration, attribute
Valuesstring, long, double/float, boolean, keyword, symbol, ref, instant (#inst), uuid (#uuid), uri, bigint/bigdec (as numbers), tuple (as vectors), fn
Queryq query, list/map/string queries; find specs rel, [?x ...], [?a ?b], ?x .; pull in :find; :with; :in scalars, tuples, collections, relations, several sources, rules (%); data patterns incl. tx and added positions; predicates and functions with all binding forms; not not-join or or-join and; rules, recursive and over cyclic data; aggregates count count-distinct sum min max avg median distinct (min n ?x) (max n ?x) sample rand; get-else get-some missing? ground tuple untuple fulltext
Pullpull pull-many; *, :db/id, reverse attributes, nested maps, recursion (... and depth limits), :as :limit :default (vector or list syntax) and legacy (limit ...) (default ...), component attributes pulled recursively
Entitiesentity touch entity-db; lazy lookup, keyword/get access, keys, reverse navigation (:ns/_attr), enums as keywords, cardinality-many as sets
Timeas-of since history (by t, tx id or instant), as-of-t since-t basis-t next-t is-history filter is-filtered
Indexesdatoms seek-datoms index-range over :eavt :aevt :avet :vaet; datoms support :e :a :v :tx :added and nth
Loglog tx-range tx-report-queue (.poll .take .peek .isEmpty .size) remove-tx-report-queue
Errorsex-info with Datomic's :db/error codes (:db.error/unique-conflict, :db.error/cas-failed, :db.error/wrong-type-for-attribute, :db.error/not-an-entity, :db.error/datoms-conflict, …); a failed transact throws when dereferenced

Durable storage

A datomic:sql URI with a JDBC URL makes a database durable, in SQLite or in PostgreSQL:

;; one machine: a SQLite file
(def uri "datomic:sql://hello?jdbc:sqlite:/var/lib/app/datomic.db")
;; any number of machines: a PostgreSQL database, from any provider
(def uri (str "datomic:sql://hello?jdbc:postgresql://db.example.com:5432/app"
              "?user=app&password=" (System/getenv "PGPASSWORD") "&sslmode=require"))

(d/create-database uri)
(def conn (d/connect uri))

The storage is created on first use and can hold any number of databases. create-database, delete-database, rename-database and get-database-names (with datomic:sql://*?<jdbc-url>) act on its catalog. A PostgreSQL URL is handed to libpq without its jdbc: prefix, so anything libpq accepts works, TLS options included.

  • What is stored. Each transaction is stored as one log row: the datoms it produced, its tempids and the id counters after it. connect replays the log to rebuild the database, so every id and value comes back as the transaction made it. The transaction logic never runs twice.
  • Several processes. Any number of processes can share a storage. transact takes the storage's write lock, applies what others have committed since it last looked, transacts, and stores the result before releasing the lock. The lock is BEGIN IMMEDIATE on SQLite and a transaction-scoped advisory lock on PostgreSQL. So each process acts as its own transactor, one at a time.
  • Two round trips per write on PostgreSQL. A write sends two multi-statement queries. The first begins, takes the lock, checks for a transactor, and reads what others committed. The second appends, notifies, and commits. Latency to the server matters more than anything else here, so keep the database in the same region as the peers.
  • Seeing other writers. d/db, sync and reading a tx-report-queue pick up other writers' transactions, with their :tempids.
  • Push. On PostgreSQL every commit sends a NOTIFY, so a process waiting for a transaction wakes as soon as it lands. That covers dereferencing a queued transaction and a tx-report-queue's .take, which blocks until the next transaction. SQLite has no way to tell other processes, so there waiting means looking again every couple of milliseconds, and .take returns nil when the queue is empty.
  • Transaction functions. A :db/fn holds a Clojure fn, which can't be stored. Installing one in a stored database fails with :db.error/not-storable.
  • Releasing. release drops a stored connection from the cache, so the next connect rebuilds it from storage.
  • Transactor. Running a transactor makes it the only process that writes the log (see Running a transactor).
  • Requirements. Storage uses clonim's clonim.sqlite and clonim.postgres. They load libsqlite3 or libpq when a stored database is first used.

Running a transactor

clonim build script/transactor.clj --source-path src -o nitomic-transactor
./nitomic-transactor 'jdbc:postgresql://host/app?user=app&password=...'
./nitomic-transactor /var/lib/app/datomic.db   # SQLite
NITOMIC_STORAGE='jdbc:postgresql://...' ./nitomic-transactor

nitomic.transactor serves every database in a storage. It claims the storage while it runs:

  • The claim. On PostgreSQL the claim is a session advisory lock, which the server drops the moment the transactor's connection goes, even if it is killed. SQLite has nothing like that, so there the transactor records a heartbeat every second instead.
  • Queued transactions. While a transactor has the storage, d/transact doesn't write the log. It queues the transaction data and returns a future at once. The transactor takes queued transactions in order. For each one, in a single transaction, it runs it, appends it to the log, and records the outcome. Dereferencing the future waits for that outcome, then returns the report or throws the transactor's error.
  • Waiting. On PostgreSQL the transactor sleeps until a peer's NOTIFY says something was queued, and peers sleep until its NOTIFY says their transaction is done.
  • Clock. The transactor stamps :db/txInstant with its own clock, as Datomic's does.
  • Only one at a time. A second transactor refuses to start (:db.error/transactor-running).
  • When it stops. Peers go back to writing the log themselves, so the databases stay writable. On PostgreSQL that happens at once; on SQLite, after three missed heartbeats. A transaction still waiting in the queue at that point is withdrawn and fails with :db.error/transactor-unavailable.
  • What can be queued. Transaction data crosses processes as EDN, with tempids and datoms encoded. A fn can't be sent and fails with :db.error/not-storable.
  • Programmatic use. nitomic.transactor/run serves a storage until its :stop? fn returns true. start, step! and stop! drive it one batch at a time.

Deploying on Modal

On Modal every container is its own machine, and containers come and go, so use PostgreSQL storage from any provider (Neon, Supabase, RDS, your own). A Modal Volume can't hold a shared SQLite file: its changes only reach other containers through commit() and reload(), which SQLite's locking knows nothing about.

  • Build. Build each program with clonim build, and copy the binaries into an image that has libpq5 and libpcre3 installed.
  • Credentials. Keep the connection URL in a Modal Secret and read it with System/getenv.
  • Direct connections. Use the provider's direct connection string, not a transaction-mode pooler (PgBouncer, Neon's -pooler host, Supabase's port 6543). A pooler hands each transaction to a different server session, which breaks LISTEN and the transactor's session lock.
  • Peers. Run your app as usual (for example behind @modal.web_server). Every container is a peer: it connects, replays the log once, and then catches up.
  • Transactor (optional). Run nitomic-transactor as a single always-on function (min_containers=1, max_containers=1). If Modal restarts it, peers write the log themselves until it is back.

The repository's deploy/modal/ directory does all of this. It has a Modal app with an HTTP API over nitomic peers and an optional transactor, deployed with modal deploy deploy/modal/app.py. With Neon as the PostgreSQL, see deploy/modal/README.md: use Neon's direct (not pooled) connection string, and deploy without the transactor (NITOMIC_TRANSACTOR=0) if the compute should scale to zero.

Differences from Datomic

  • Storage is memory, SQLite or PostgreSQL. datomic:sql URIs with a jdbc:sqlite: or jdbc:postgresql: URL keep databases durably (see Durable storage). Every other URI protocol (mem, dev, ddb, …) names an in-process database that is gone when the process exits. The transactor is optional: without one, each process takes the storage's write lock to transact.
  • No runtime code compilation. Database functions can't be Clojure source strings. :db/fn holds a Clojure fn, which d/function passes through. Query functions are found in a built-in table of clojure.core and string functions (including .compareTo, .startsWith, and the like). You can also pass them as inputs or register them with d/register-fn!. The JVM would resolve any qualified symbol instead. See test/native.clj.
  • Fulltext tokenizes on letters and digits, lower-cases, drops English stop words, and matches any query term (with term* prefixes). That matches the default Lucene analyzer on typical text. It is not Lucene's full query syntax, and scores are all 1.0.
  • Not implemented: composite tuple attributes (:db/tupleAttrs), :db/ensure and entity specs, excision, d/index-pull, and database stats beyond counts.
  • Printing. Query results are Clojure sets and vectors rather than Java collections. An entity prints as {:db/id n}, and a touched entity prints as its attribute map.

How it is tested

nitomic is tested by comparison with Datomic itself. Each test program uses only datomic.api, runs unchanged on the JVM against Datomic Pro and natively against nitomic, and prints its results in a canonical order. The native output must match the recorded JVM output line for line.

The test programs

programwhat it coversexpected output
examples/seattle/getting_started.cljDatomic's getting-started walkthrough over the Seattle data (the Getting started part of this book)recorded on Datomic Pro
test/features.cljthe rest of the API, including error casesrecorded on Datomic Pro
test/native.cljnitomic-only extensions, such as d/register-fn!written for nitomic
test/storage.cljdurable storage (SQLite, or PostgreSQL via NITOMIC_TEST_STORAGE): replay, two connections sharing a storage, the catalogwritten for nitomic
test/transactor.cljthe transactor: queued transactions, errors, tempids, stopping and withdrawalwritten for nitomic

Running the tests

CLONIM=path/to/clonim/bin/clonim test/run.sh
# or, with clonim on PATH
nimble test

test/run.sh runs each program with clonim run ... --source-path src, diffs its output against test/expected/<name>.out, and prints ok or FAIL with the first lines of the diff:

ok   getting_started
ok   features
ok   native
ok   storage
ok   transactor

The storage and transactor tests use a SQLite file unless NITOMIC_TEST_STORAGE names another storage. Given a PostgreSQL URL, they must print exactly the same output.

CI (.github/workflows/test.yml) builds clonim with Nim 2.2.12 and runs the same script on every push and pull request twice: once as is, and once against a PostgreSQL service container.

Re-recording the expected output

The expected outputs of the first two programs were recorded by script/reference.sh on the JVM against Datomic Pro 1.0.7705. To re-record them from a distribution:

curl -O https://datomic-pro-downloads.s3.amazonaws.com/1.0.7705/datomic-pro-1.0.7705.zip
unzip datomic-pro-1.0.7705.zip
script/reference.sh datomic-pro-1.0.7705

The script runs each program with clojure.main on the peer jar, and drops the JVM's log lines and reflection warnings so that only program output is kept.

Writing a comparable program

To make a new program comparable across the two platforms, follow the walkthrough's conventions:

  • print results through a canonical printer (see the show helpers in the Overview), so that set order doesn't matter;
  • print instants as a placeholder and leave tempids out, since they differ between runs and platforms;
  • end with (System/exit 0) so the JVM run terminates.

Source layout

filewhat it does
src/datomic/api.cljthe public API
src/datomic/db.cljid-literal, the #db/id reader function
src/nitomic/db.cljdatabase values: covering indexes as nested persistent maps, the schema cache, as-of/since/history views
src/nitomic/tx.cljtransactions: expansion, tempid resolution and upsert, id allocation, index updates, schema installation
src/nitomic/query.cljDatalog, rules and aggregates
src/nitomic/pull.cljthe pull API
src/nitomic/entity.cljlazy entities (a deftype over ILookup/Seqable)
src/nitomic/storage.cljdurable storage in SQLite or PostgreSQL: the catalog, transaction log and queue, replay, locks and notifications
src/nitomic/transactor.cljthe transactor: serves a storage's queue and writes its log
src/nitomic/types.cljdatoms and tempids
src/nitomic/bootstrap.cljDatomic's bootstrap datoms

A database is an immutable map, so every database value stays valid. Its covering indexes ({e {a {v tx}}}, {a {e {v tx}}}, {a {v {e tx}}} and {v {a {e tx}}} for refs) make every bound prefix a hash lookup. The sorted orders that d/datoms promises are produced on demand. As in Datomic, query clauses run in the order written.