Skip to content

Project starter #1

Description

@thorwhalen

Previous (old) issue: i2mint/py2store#40 on the subject.

We'll restrict ourselves, for now, to a following types of data (in order of priority):

  • fixed-rate fvs (i.e. fixed size numerical vectors).
  • fixed-rate snips (i.e. uni-dimensional sequences of non-negative integers (so can use uint8, unint16, unint32...))
  • multi-dimensional non-regular numerical timeseries (here, should use one channel to represent timestamps)

These cover most of the (pragmatic ground), but it would be useful to have a few data converters to be able to easily cast data to the forms accepted by the serializers. For this we need to (1) decide what our "normal" input format is and (2) set things up so we can add a growing number of types we'll translate from. For the normal input format I'd say the first argument is always an iterable of numbers or fixed-size sequences (tuples or lists say) of numbers. Additionally, there'll be some optional arguments to describe things about the data (to avoid having to peek) and other meta data (that might not be the data's concern, but is the serializers.

Serializers should go from (data, [meta]) to bytes, and deserializers from bytes to (data, [meta]).

The persistence (saving the bytes) is a separate concern.

Within the persistence concern is a concern of indexing -- namely according to time. Here, we need to tackle aspects such as block storage

  • dividing into pieces, and bringing pieces together to recreate the whole
  • external and internal indexing to be able to satisfy interval queries (give me the data between bt and tt timestamps)

Note: When timeseries are created by fixed-step chunkers, the data rate is determined by the step_step (NOT the step_size).

Session-block storage

data_provider = get_timeseries(src, bt, tt)

# alternative:
my_data = mk_timeseries_interval_accessor(src, ...)
data_provider = my_data[bt:tt]

# and then use...

Resources to check out

Would be good to find some resources on codec design and formats. Something describing a principled approach so we don't have to rethink the wheel, and we can use the words the community uses to describe things (such as "codec", "container", "header", etc). Here are a few things:

It seems IFF/TLV could be used with for headers. But it seems wasteful to use this for our chunks/frames, since we are only dealing with fixed structure in our case.

Builtins

Might want to check this codec module out.

And for sure, we'll need the struct module.

The following probably use the above (not sure though):

wave

chunk

More like this in this list

Third party

soundfile

our stuff

Peek at an iterator -- useful when you want to see what format it's elements have (without consuming it) --

Somewhat related modules of ours:
https://github.com/otosense/hear/blob/3a87757d3fd094c8834c6f10e845d5a45592e026/hear/regular_panel_data.py
https://github.com/i2mint/py2store/blob/813bb853a28be1eef6454b9d7d9be5ddb1f9b7b1/py2store/utils/affine_conversion.py
https://github.com/i2mint/py2store/blob/6d525784c9212a4839c42dfbc4fb9d427b959811/py2store/utils/timeseries_caching.py

https://github.com/otosense/hear/blob/3a87757d3fd094c8834c6f10e845d5a45592e026/hear/session_block_stores.py
https://github.com/otosense/hear/blob/3a87757d3fd094c8834c6f10e845d5a45592e026/hear/stores.py

┆Issue is synchronized with this Asana task by Unito

Activity

  1. owen7lloyd commented on Jun 18, 2021

    @owen7lloyd
    Collaborator

    Notes on things to do now:

    • struct.iter_unpack - does it do anything for us (might chunk things for us)
    • comments / uncle bob refactoring
    • look at codec, what is it about, how can we use it

    Looking forward:

    • replicate the main functionality of soundfile (i.e. read and write of audio)
    • add struct to the dicts in util
    • find the name of this (fixed size of chunks with fixed size of bytes)
    • try looking at raw video format
    • implicitly make functionality easier for user:
      • use peek to make the interface easier (find the number of channels, format, etc.)
        • function takes first sample and decides which format to apply (user can override)
    • deal with rows of a dataframe (list of dicts)
    • serialize the meta into a header
    • express the meta in a normalized dict fashion and pickle it, put the pickle as the header
      use type and length to determine length of header (pickle varies in length) to deserialize
  2. thorwhalen commented on Aug 17, 2026

    @thorwhalen
    MemberAuthor

    Verified against current code (0.1.38) — recommend keep-open-and-narrow

    Five years on, most of this issue's core shipped and two of its asks did not. Item by item:

    SHIPPED

    • Fixed-rate fvs — mk_codec(chk_format, n_channels, chk_size_bytes) (base.py:35) building StructCodecSpecs + ChunkedEncoder/ChunkedDecoder/IterativeDecoder. Verified: mk_codec('fff') round-trips [[1,2,3],[4,5,6]].
    • Fixed-rate snips — same path with struct's unsigned chars; mk_codec('B')/('H')/('I') each round-trip [0,1,255]. But no unsigned format appears in any test parametrization (the suite exercises only 'h', 'd' and byte-order prefixes), and "unsigned"/"uint" appears nowhere in the README. Mechanism shipped, use case undocumented and untested.
    • The "normal" input format — codified verbatim as the issue proposed, at base.py:17-19: Sample = Any, Frame = Union[Sample, Sequence[Sample]], Frames = Iterable[Frame]. The "optional arguments to avoid having to peek" landed as n_channels / chk_size_bytes.
    • Research vocabulary — struct is the engine, wave powers the header codec, soundfile's naming is absorbed into num_type_synonyms (audio.py). IFF/TLV were considered and deliberately not adopted.

    STILL OPEN — and this is the main unbuilt piece

    • Data converters, "set things up so we can add a growing number of types we'll translate from" — the translation table is closed and string-keyed: util.py:28 is type_to_struct = {"<class 'int'>": 'h', "<class 'float'>": 'd'}, consumed by get_struct via str(type(x)). No registration hook, no dispatch — the opposite of the open-closed design the issue asked for. Two live consequences:

      • specs_from_frames(np.array([1., 2., 3.], dtype='float32')) → AttributeError: Unknown data format
      • specs_from_frames([1, 2, 100000]) cheerfully infers 'h', then dies at encode time with struct.error: 'h' format requires -32768 <= number <= 32767

      That is the single most user-visible gap in the package.

    PARTIALLY SHIPPED

    • Serializers go from (data, [meta]) to bytes — fully worked in the audio module, thin in the generic one. MetaEncoder/MetaDecoder write only the column names into the header (pack('h', len(s)) + '.'-joined names, base.py:276), so a blob carries names but not dtype/chk_format/rate — it cannot be decoded without out-of-band knowledge. The WAV path, by contrast, is self-describing. The follow-up comment on this issue ("serialize the meta into a header… express the meta in a normalized dict fashion") names exactly this gap.

    STILL OPEN, but nothing is timestamp-aware

    • Multi-dimensional non-regular timeseries with a timestamp channel — expressible only because a timestamp is just another struct channel (mk_codec('qhh') round-trips [(1000,3,-1),(1017,4,-2)]). Grepping timestamp|timeseries|irregular|non-regular|sample_rate across base.py and audio.py returns one hit — the unused ByteChunker alias — and zero in the README. No API marks a channel as time.

    OBSOLETE for this repo

    • Persistence: block storage, chunk/reassemble, external+internal time indexing, get_timerange interval queries — recode contains no storage code at all (4 files, ~1250 lines, no store, no bt/tt, no slicing). The issue scoped this out itself ("the persistence is a separate concern") and setup.cfg scopes the package to serialization. This belongs to the store packages.

    Recommendation

    Rewrite the body down to two items and retitle:

    1. An open/extensible type → struct-format converter registry, replacing the string-keyed dict, with range-aware inference (the [1, 2, 100000] case above as its acceptance test).
    2. A self-describing header/meta codec for the generic path, with the timestamp-channel timeseries as its first customer.

    Delete the persistence section (out of scope, and the issue says so) and the "resources to check out" section (it has served its purpose — struct, wave, soundfile naming and the iterator-peek trick are all consumed in the code).

    Leaving it as a 2021 project-starter means the two real gaps stay invisible inside a mostly-done checklist.

    Established by a verified audit alongside the #4 fix (#11). Not acting on it — the rewrite is a maintainer call.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentation

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions