Initialisation
Often advanced use case require reference data to support the processing of calculations or business logic. Joule enables this by providing data priming at initialisation and enrichment processing
The initialisation process imports at startup data from the local file store into the in-memory SQL database. This data is intended be used This is an optional feature.
Data is accessibility points provided
Select projection
In-memory SQL API
Example
The below example loads two data files in to independent in-memory SQL database tables using a CSV and parquet files. The CSV file contains nasdaq company information which will can be consisted static and therefore will reside within the reference_data schema whereas the a parquet file primes the metrics engine with pre-calculated metrics.
Data Import
Joule has an embedded SQL engine that when enabled various features such as metrics, event capturing, exporting and reference data can be taken advantage of.
Data imported in this manner is typically for static reference data which supports key processing functions within the event stream pipeline.
Attribute | Description | Data Type | Required |
---|---|---|---|
schema | Global database schema when set can used for any import definition where schema is not defined. Default schema reference_data | String | |
csv | List of CSV data import configurations | See CSV attributes | |
parquet | List of parquet data import configurations | Seee parquet attributes |
CSV and Parquet common DLS elements
There are common keywords used in both parquet and CSV importing methods
Attribute | Description | Data Type | Required |
---|---|---|---|
schema | Database schema to create and appy table import function | String | |
table | Target table to import data into | String | |
drop table | Drop existing table before import. This will cause a table recreation. | Boolean Default true | |
index | Create an index on the created table | See Below |
Index
If this optional field is supplied the index is recreated once the data has been imported.
Attribute | Description | Data Type | Required |
---|---|---|---|
fields | A list of table fields to base index on. | String | |
unique | True for a unique index | Boolean Default true |
Parquet Import
Parquet formatted files can be imported into the system. Note an index cannot be created over a view.
Example
Attribute | Description | Data Type | Required |
---|---|---|---|
asView | Create a view over the files'. This will mean disk IO for every table query. Consider | Boolean Default false | |
files | List of files of the same type to be imported | String list |
CSV
Data can be imported from CSV files using a supported set of delimiters. The key difference between parquet and CSV is you can control the table definition. Joule by default will try to create a target table based upon a sample set of data assuming a header exists on the first row.
Example
Attribute | Description | Data Type | Required |
---|---|---|---|
table definition | Custom SQL table definition used when provided. This will override the | String | |
file | Name and path of the file to import. | String | |
delimiter | Field delimiter to use. Supported delimiters include | String Default | | |
date format | User specified date format. | String Default: System | |
timestamp formate | User specified timestamp format. | String Default: System | |
sample size | Number of rows to use to determine types and table structure. | Integer Default: 1024 | |
skip | Number of rows to skip when auto generating a table using a sample size. | Integer Default: 1 | |
header | Flag to indicate the first line in the file contains a header. | Boolean Default: true | |
auto detect | Auto detect table format by taking a sample size of data. If set to | Boolean Default: true |
Application example
The below example primes the process with reference and metrics data which is used for event enrichment processing.
Access
To access data loaded using this method Joule provides a SQL API. See further documentation on how to use this in your custom components
Last updated