Repository navigation
Project starter #1
Description
Activity
- addeddocumentationImprovements or additions to documentationImprovements or additions to documentation
on Jun 18, 2021 Notes on things to do now:
struct.iter_unpack- does it do anything for us (might chunk things for us)- comments / uncle bob refactoring
- look at
codec, what is it about, how can we use it
Looking forward:
- replicate the main functionality of
soundfile(i.e. read and write of audio) - add struct to the dicts in util
- find the name of this (fixed size of chunks with fixed size of bytes)
- try looking at raw video format
- implicitly make functionality easier for user:
- use peek to make the interface easier (find the number of channels, format, etc.)
- function takes first sample and decides which format to apply (user can override)
- use peek to make the interface easier (find the number of channels, format, etc.)
- deal with rows of a dataframe (list of dicts)
- serialize the meta into a header
- express the meta in a normalized dict fashion and pickle it, put the pickle as the header
use type and length to determine length of header (pickle varies in length) to deserialize
Verified against current code (0.1.38) — recommend keep-open-and-narrow
Five years on, most of this issue's core shipped and two of its asks did not. Item by item:
SHIPPED
- Fixed-rate fvs —
mk_codec(chk_format, n_channels, chk_size_bytes)(base.py:35) buildingStructCodecSpecs+ChunkedEncoder/ChunkedDecoder/IterativeDecoder. Verified:mk_codec('fff')round-trips[[1,2,3],[4,5,6]]. - Fixed-rate snips — same path with struct's unsigned chars;
mk_codec('B')/('H')/('I')each round-trip[0,1,255]. But no unsigned format appears in any test parametrization (the suite exercises only'h','d'and byte-order prefixes), and "unsigned"/"uint" appears nowhere in the README. Mechanism shipped, use case undocumented and untested. - The "normal" input format — codified verbatim as the issue proposed, at
base.py:17-19:Sample = Any,Frame = Union[Sample, Sequence[Sample]],Frames = Iterable[Frame]. The "optional arguments to avoid having to peek" landed asn_channels/chk_size_bytes. - Research vocabulary —
structis the engine,wavepowers the header codec, soundfile's naming is absorbed intonum_type_synonyms(audio.py). IFF/TLV were considered and deliberately not adopted.
STILL OPEN — and this is the main unbuilt piece
-
Data converters, "set things up so we can add a growing number of types we'll translate from" — the translation table is closed and string-keyed:
util.py:28istype_to_struct = {"<class 'int'>": 'h', "<class 'float'>": 'd'}, consumed byget_structviastr(type(x)). No registration hook, no dispatch — the opposite of the open-closed design the issue asked for. Two live consequences:specs_from_frames(np.array([1., 2., 3.], dtype='float32'))→AttributeError: Unknown data formatspecs_from_frames([1, 2, 100000])cheerfully infers'h', then dies at encode time withstruct.error: 'h' format requires -32768 <= number <= 32767
That is the single most user-visible gap in the package.
PARTIALLY SHIPPED
- Serializers go from (data, [meta]) to bytes — fully worked in the audio module, thin in the generic one.
MetaEncoder/MetaDecoderwrite only the column names into the header (pack('h', len(s))+'.'-joined names,base.py:276), so a blob carries names but not dtype/chk_format/rate — it cannot be decoded without out-of-band knowledge. The WAV path, by contrast, is self-describing. The follow-up comment on this issue ("serialize the meta into a header… express the meta in a normalized dict fashion") names exactly this gap.
STILL OPEN, but nothing is timestamp-aware
- Multi-dimensional non-regular timeseries with a timestamp channel — expressible only because a timestamp is just another struct channel (
mk_codec('qhh')round-trips[(1000,3,-1),(1017,4,-2)]). Greppingtimestamp|timeseries|irregular|non-regular|sample_rateacrossbase.pyandaudio.pyreturns one hit — the unusedByteChunkeralias — and zero in the README. No API marks a channel as time.
OBSOLETE for this repo
- Persistence: block storage, chunk/reassemble, external+internal time indexing,
get_timerangeinterval queries —recodecontains no storage code at all (4 files, ~1250 lines, no store, nobt/tt, no slicing). The issue scoped this out itself ("the persistence is a separate concern") andsetup.cfgscopes the package to serialization. This belongs to the store packages.
Recommendation
Rewrite the body down to two items and retitle:
- An open/extensible type → struct-format converter registry, replacing the string-keyed dict, with range-aware inference (the
[1, 2, 100000]case above as its acceptance test). - A self-describing header/meta codec for the generic path, with the timestamp-channel timeseries as its first customer.
Delete the persistence section (out of scope, and the issue says so) and the "resources to check out" section (it has served its purpose —
struct,wave,soundfilenaming and the iterator-peek trick are all consumed in the code).Leaving it as a 2021 project-starter means the two real gaps stay invisible inside a mostly-done checklist.
Established by a verified audit alongside the #4 fix (#11). Not acting on it — the rewrite is a maintainer call.
- Fixed-rate fvs —
Previous (old) issue: i2mint/py2store#40 on the subject.
We'll restrict ourselves, for now, to a following types of data (in order of priority):
These cover most of the (pragmatic ground), but it would be useful to have a few data converters to be able to easily cast data to the forms accepted by the serializers. For this we need to (1) decide what our "normal" input format is and (2) set things up so we can add a growing number of types we'll translate from. For the normal input format I'd say the first argument is always an iterable of numbers or fixed-size sequences (tuples or lists say) of numbers. Additionally, there'll be some optional arguments to describe things about the data (to avoid having to peek) and other meta data (that might not be the data's concern, but is the serializers.
Serializers should go from (data, [meta]) to bytes, and deserializers from bytes to (data, [meta]).
The persistence (saving the bytes) is a separate concern.
Within the persistence concern is a concern of indexing -- namely according to time. Here, we need to tackle aspects such as block storage
btandtttimestamps)Note: When timeseries are created by fixed-step chunkers, the data rate is determined by the
step_step(NOT thestep_size).Session-block storage
Resources to check out
Would be good to find some resources on codec design and formats. Something describing a principled approach so we don't have to rethink the wheel, and we can use the words the community uses to describe things (such as "codec", "container", "header", etc). Here are a few things:
It seems IFF/TLV could be used with for headers. But it seems wasteful to use this for our chunks/frames, since we are only dealing with fixed structure in our case.
Builtins
Might want to check this codec module out.
And for sure, we'll need the struct module.
The following probably use the above (not sure though):
wave
chunk
More like this in this list
Third party
soundfile
our stuff
Peek at an iterator -- useful when you want to see what format it's elements have (without consuming it) --
Somewhat related modules of ours:
https://github.com/otosense/hear/blob/3a87757d3fd094c8834c6f10e845d5a45592e026/hear/regular_panel_data.py
https://github.com/i2mint/py2store/blob/813bb853a28be1eef6454b9d7d9be5ddb1f9b7b1/py2store/utils/affine_conversion.py
https://github.com/i2mint/py2store/blob/6d525784c9212a4839c42dfbc4fb9d427b959811/py2store/utils/timeseries_caching.py
https://github.com/otosense/hear/blob/3a87757d3fd094c8834c6f10e845d5a45592e026/hear/session_block_stores.py
https://github.com/otosense/hear/blob/3a87757d3fd094c8834c6f10e845d5a45592e026/hear/stores.py
┆Issue is synchronized with this Asana task by Unito