ParallelLoader
ParallelLoader Objects
class ParallelLoader(ParallelQuery.ParallelQuery)
Parallel and Batch Loader for ApertureDB
This takes a dataset (which is a collection of homogeneous objects) or a derived class, and optimally inserts them into database by splitting them into batches, and passing the batches to multiple workers.
Hierarchy of Data Loaders:
It accepts any Subscriptable that yields (commands, blobs) (e.g., standard data loaders), and the dataset is passed to ingest() rather than the constructor. Examples of supported loaders include:
aperturedb.BBoxDataCSV.BBoxDataCSVaperturedb.BlobDataCSV.BlobDataCSVaperturedb.ConnectionDataCSV.ConnectionDataCSVaperturedb.DescriptorDataCSV.DescriptorDataCSVaperturedb.DescriptorSetDataCSV.DescriptorSetDataCSVaperturedb.EntityDataCSV.EntityDataCSVaperturedb.ImageDataCSV.ImageDataCSVaperturedb.PolygonDataCSV.PolygonDataCSVaperturedb.VideoDataCSV.VideoDataCSV- A class derived from
aperturedb.PyTorchData.PyTorchData - A class derived from
aperturedb.KaggleData.KaggleData
get_existing_indices
def get_existing_indices() -> dict
Returns the existing ordered indexes, as used for constraint lookups.
Returns:
dict- A dictionary of the form{"entity": {"class_name": {"property_name"}}}, with a"connection"key for connection indexes.
query_setup
def query_setup(generator: Subscriptable) -> None
Runs the setup for the loader, which includes creating indices. Currently, it only creates indices for the properties that are used for constraint.
Will only run when the argument generator has a get_indices method that returns a dictionary of the form:
{
"entity": {
"class_name": ["property_name"]
},
}
or
{
"connection": {
"class_name": ["property_name"]
},
}
Arguments:
generatorSubscriptable - The Subscriptable object that is being ingested
ingest
def ingest(generator: Subscriptable,
batchsize: int = 1,
numthreads: int = 4,
stats: bool = False,
transformers: list = None) -> None
Method to ingest data into the database
Arguments:
generatorSubscriptable - The list of data, or a class derived from Subscriptable to be ingested.batchsizeint, optional - The size of batch to be used. Defaults to 1.numthreadsint, optional - Number of workers to create. Defaults to 4.statsbool, optional - If stats need to be presented, realtime. Defaults to False.transformerslist, optional - A Transformer class, a callable, or a list of Transformer classes/callables to apply to the data. Defaults to None.