!!! abstract “”
How to use `open_data_store` in read, write, and append modes with directory, zip, and SQLite backends, iterate over members, and inspect `.completed`, `.not_completed`, and `.summary_<methods>`.
A data store is just a “container”. To open a data store you use the
open_data_store() function. To load the data for a member
of a data store you need an appropriately selected loader type of
app.
All data store classes can be iterated over, indexed, checked for
membership. These operations return a DataMember object. In
addition to providing access to members, the data store classes have
convenience methods for describing their contents and providing
summaries of log files that are included and of the
NotCompleted members (see not completed).
Use the open_data_store() function, illustrated below.
Use the mode argument to identify whether to open as read only
(mode="r"), write (mode="w") or
append(mode="a").
We open the zipped directory described above, defining the filenames
ending in .fa as the data store members. All files within
the directory become members of the data store (unless we use the
limit argument).
```python { linenums=“1” notest } from scinexus import open_data_store
dstore = open_data_store(“data/raw.zip”, suffix=“fa”, mode=“r”) # (1)! print(dstore)
dstore.describe # (2)!
m = dstore[0] # (3)!
for m in dstore[:5]: # (4)! print(m)
m.read()[:20] # (5)!
1. Open a data store.
2. The `.describe` property summarises the contents.
3. You can index like any Python sequence.
4. Or loop over members.
5. And read data from a member.
<!-- [[[end]]] -->
!!! note
For a `DataStoreSqlite` member, the default data storage format is bytes. So reading the content of an individual record is best done using the `load_db` app.
### Making a writeable data store
The creation of a writeable data store is specified with `mode="w"`, or (to append) `mode="a"`. In the former case, any existing records are overwritten. In the latter case, existing records are ignored.
In a directory store, the `suffix` you open with names every *completed* record it writes. The identifier you pass to `write()` supplies the stem only, so `"brca1"`, `"brca1.fa"` and `"brca1.genbank"` all become `brca1.fasta` in a store opened with `suffix="fasta"`. Not-completed records, logs and checksums are kept in their own subdirectories under their own extensions, and the store's suffix does not apply to them.
!!! warning "A compression suffix is refused, not replaced"
Compression is the one part of the name that says how to read the record back, so a directory store will not quietly swap it. If the identifier names a compression the store does not write, `write()` raises `ValueError` rather than storing the record under a different name.
This part is specific to directory stores, since a SQLite store has no suffix to contradict. The rule that an identifier must name a record — a non-blank, non-hidden stem — applies to both.
```python { notest }
dstore = open_data_store("results", suffix="fasta", mode="w")
dstore.write(unique_id="brca1.fasta.gz", data=seqs)
# ValueError: identifier 'brca1.fasta.gz' names .gz, but a record
# stored as .fasta carries no compression
```
Open the store with the compression in its suffix instead, and its completed records are gzipped:
```python { notest }
dstore = open_data_store("results", suffix="fasta.gz", mode="w")
dstore.write(unique_id="brca1", data=seqs) # brca1.fasta.gz, gzipped
```
Not-completed records are always plain `.json`, whatever the store's own suffix is, so a compressed identifier is refused there too — including in a `suffix="fasta.gz"` store, where `write()` accepts `"brca1.fasta.gz"` and `write_not_completed()` does not.
## `DataStoreSqlite` stores serialised data
When you specify a Sqlitedb data store as your output (by using `open_data_store()`) you write multiple records into a single file making distribution easier.
!!! warning
The process which creates a Sqlitedb "locks" the file. If that process exits unnaturally (e.g. the run that was producing it was interrupted) then the file may remain in a locked state. If the db is in this state, `scinexus` will not modify it unless you explicitly unlock it.
### Closing a Sqlitedb data store
The lock is released by `close()`, and only by `close()`. That is what gives a lock you find on a file its meaning: it says the session that took it did not finish. So call `close()` once you have finished with a writable store, after reading whatever you need from it.
```python { notest }
out_dstore = open_data_store("results.sqlitedb", mode="w")
# ... write to it, then read your summaries from it ...
out_dstore.close()
A store that is garbage collected without being closed warns you and names the file, because it has left a lock behind that the next run will refuse to write over. Closing a data store ends access to it – trying to read from one afterwards raises an exception rather than returning stale answers.
Directory and zip data stores take no such lock, and closing one does
nothing. close() is defined on every data store so that
code holding whatever open_data_store() returned can close
it without first asking which kind it got.
This is represented in the display as shown below.
python { linenums="1" notest } {"completed": 175, "not_completed": 0, "logs": 1, "title": "Unlocked db store."}
To unlock, open the store in a writable mode and execute the following:
python { notest } dstore = open_data_store("data/demo-locked.sqlitedb", mode="a") dstore.unlock(force=True)
The store has to be writable, since unlocking writes to it (a read
only store ignores the call). force=True is needed when the
lock was taken by another process, which is the usual case for a store
left behind by an interrupted run. Overriding it is meant to be a
deliberate act, because the records in such a store were never confirmed
complete.
If you use the apply_to() method, a
scitrack logfile will be stored in the data store. This
includes useful information regarding the run conditions that produced
the contents of the data store.
python { linenums="1" notest } # [{'time': '2019-07-24 14:42:56', 'name': 'logs/load_unaligned-progressive_align- # write_db-pid8650.log', 'python_version': '3.7.3', 'who': 'gavin', 'command': # '/Users/gavin/miniconda3/envs/c3dev/lib/python3.7/site- # packages/ipykernel_launcher.py -f [...]
Log files can be accessed via a special attribute.
python { linenums="1" notest } # [DataMember(data_store=/Users/gavin/repos/SciNexus/docs/data/demo- # locked.sqlitedb, unique_id=logs/load_unaligned-progressive_align-write_db- # pid8650.log)]
Each element in that list is a DataMember which you can
use to get the data contents. The following
python { notest } print(dstore.logs[0].read()[:225])
Produces
python { linenums="1" notest } # 2019-07-24 14:42:56 Eratosthenes.local:8650 INFO system_details : # system=Darwin Kernel Version 18.6.0: Thu Apr 25 23:16:27 PDT 2019; # root:xnu-4903.261.4~2/RELEASE_X86_64 2019-07-24 14:42:56 # Eratosthenes.local:8650 INFO python
When apps declare citations, those citations are automatically saved
alongside your results when you use apply_to().
```python { linenums=“1” notest } import pathlib import shutil from citeable import Software from scinexus import define_app, open_data_store from cogent3 import get_app from cogent3.app.typing import AlignedSeqsType
my_cite = Software( author=[“Doe, J”], title=“My Sequence Filter”, year=2025, )
@define_app(cite=my_cite) def strict_filter(val: AlignedSeqsType) -> AlignedSeqsType: return val.omit_bad_seqs()
in_dstore = open_data_store(“data/raw.zip”, suffix=“fa”, limit=5) out_dstore = open_data_store(“cited_results”, suffix=“fa”, mode=“w”)
loader = get_app(“load_aligned”, moltype=“dna”, format_name=“fasta”) writer = get_app(“write_seqs”, data_store=out_dstore, format_name=“fasta”) process = loader + strict_filter() + writer result = process.apply_to(in_dstore) result.write_bib(“my_analysis.bib”) print(pathlib.Path(“my_analysis.bib”).read_text())
```
The summary_citations property returns a table of all
citations stored in the data store (line 24). Export to BibTeX with
write_bib() (line 26).
!!! note ReadOnlyDataStoreZipped supports reading stored
citations but not writing them.