@jbr-experiment/ldbc-snb
v6.2.1
Published
LDBC Social Network Benchmark (Interactive) experiment handler for JBR
Maintainers
Readme
JBR Experiment - LDBC SNB
A jbr experiment type for the LDBC Social Network Benchmark (SNB) Interactive workload, executed over a single (centralized) RDF dataset.
The dataset is generated in Turtle using the LDBC SNB Hadoop datagen, merged into a single N-Triples file, and queried using SPARQL versions of the Interactive short (IS1-7) and complex (IC1-14) queries.
Requirements
- Node.js (18 or higher)
- Docker (required for invoking the LDBC SNB datagen)
- jbr (required for initializing, preparing, and running experiments on the command line)
Quick start
1. Install jbr
jbr is a command line tool that enables experiments to be initialized, prepared, and started. It can be installed from the npm registry:
$ npm install -g jbror
$ yarn global add jbr2. Initialize a new experiment
Using the jbr CLI tool, initialize a new experiment:
$ jbr init ldbc-snb my-experiment
$ cd my-experimentThis will create a new my-experiment directory with default configs for this experiment type.
3. Configure the required hooks
This experiment type requires you to configure a certain SPARQL endpoint to send queries to for the hookSparqlEndpoint.
A value for this hook can be set as follows, such as sparql-endpoint-comunica:
$ jbr set-hook hookSparqlEndpoint sparql-endpoint-comunica4. Prepare the experiment
In order to run all preprocessing steps, such as creating all required datasets, invoke the prepare step:
$ jbr prepareThis performs the following steps, where each step is skipped if its output already exists (unless jbr prepare -f is used):
- Run the LDBC SNB datagen for the configured scale factor, which produces Turtle files and substitution parameters in
generated/out-snb/. This is also skipped ifgenerated/out-snb/was removed but all files derived from it (steps 2-4) exist, such as when using pre-generated assets. - Merge all Turtle files into a single
generated/dataset.ntfile. Blank node labels of the datagen are globally unique, and are preserved. - Create
generated/parameters-persons.csv(persons frominteractive_1_param.txt) andgenerated/parameters-messages.csv(a seeded random sample of posts and comments) for the short queries. - Instantiate all query templates into
generated/queries/. - Optionally convert the dataset to HDT.
All prepared files will be contained in the generated/ directory:
generated/
dataset.hdt # If generateHdt is enabled
dataset.hdt.index.v1-1 # If generateHdt is enabled
dataset.nt
out-snb/
social_network/
social_network_activity_0_0.ttl
social_network_person_0_0.ttl
social_network_static_0_0.ttl
substitution_parameters/
interactive_1_param.txt
...
parameters-messages.csv
parameters-persons.csv
params.ini
queries/
interactive-complex-1.sparql
...
interactive-short-7.sparql5. Run the experiment
Once the experiment has been fully configured and prepared, you can run it:
$ jbr runOnce the run step completes, results will be present in the output/ directory.
Output
The following output is generated after an experiment has run.
output/query-times.csv (one line per query instantiation, aggregated over the replication rounds):
name;id;error;errorDescription;failures;hash;httpRequests;httpRequestsMax;httpRequestsMin;httpRequestsStd;replication;results;resultsMax;resultsMin;time;timeMax;timeMin;times;timestamps;timestampsMax;timestampsMin;timestampsStd;timeStd;timestampsAll
interactive-short-4;0;false;;0;2e21f854f7ee6112613cc5912d141bdb;0;0;0;0;3;1;1;1;24.666666666666668;28;22;28 24 22;21.666666666666668;24;20;1.699673171197595;2.494438257849294;"[[20],[24],[21]]"
interactive-short-4;1;false;;0;29533521492a24c11be9676db0f16540;0;0;0;0;3;1;1;1;17.666666666666668;23;13;17 23 13;17.666666666666668;23;13;4.109609335312651;4.109609335312651;"[[17],[23],[13]]"Next to this, output/query-results-raw.json contains the raw query results,
and output/logs/ contains the datagen logs and load-time.csv, the time in milliseconds between starting the endpoint and starting the measured queries (this includes the warmup rounds).
Queries
The query templates in lib/templates/queries/
are based on the SPARQL queries of SolidBench,
which in turn are based on the SNB Interactive SPARQL implementations,
using the original http://www.ldbc.eu/ldbc_socialnet/1.0/ vocabulary emitted by the datagen.
Since SPARQL has no shortest path operator, interactive-complex-13 (single shortest path) only considers paths of at most 4 hops (returning -1 otherwise),
and interactive-complex-14 (trusted connection paths) only considers shortest paths of at most 3 hops.
Short queries take a person or message IRI from the generated CSV files.
Complex queries take their parameters from the datagen's interactive_N_param.txt files.
All parameter selections are shuffled deterministically based on querySeed.
Configuration
The default generated configuration file (jbr-experiment.json) for this experiment looks as follows:
{
"@context": [
"https://linkedsoftwaredependencies.org/bundles/npm/jbr/^6.0.0/components/context.jsonld",
"https://linkedsoftwaredependencies.org/bundles/npm/@jbr-experiment/ldbc-snb/^6.0.0/components/context.jsonld"
],
"@id": "urn:jbr:my-experiment",
"@type": "ExperimentLdbcSnb",
"scale": "0.1",
"hadoopMemory": "4G",
"queryCount": 5,
"querySeed": 12345,
"generateHdt": false,
"endpointUrl": "http://localhost:3001/sparql",
"endpointUrlExternal": "http://localhost:3001/",
"queryRunnerReplication": 3,
"queryRunnerWarmupRounds": 1,
"queryRunnerRequestDelay": 0,
"queryRunnerEndpointAvailabilityCheckTimeout": 1000,
"queryRunnerUrlParams": {},
"hookSparqlEndpoint": {
"@id": "urn:jbr:my-experiment:hookSparqlEndpoint",
"@type": "HookNonConfigured"
}
}Any config changes require re-running the prepare step.
Configuration fields
scale: The SNB scale factor, such as0.1,0.3,1,3,10, ... Defaults to0.1.hadoopMemory: The maximum heap size of the Hadoop datagen, such as4G. Higher scale factors require more memory.queryCount: Number of instantiations per query template. This can not exceed the number of rows in the datagen's substitution parameter files.querySeed: Random seed for selecting query parameters.generateHdt: If adataset.hdtshould also be generated.endpointUrl: URL through which the SPARQL endpoint of thehookSparqlEndpointhook will be exposed.endpointUrlExternal: URL through which the SPARQL endpoint of thehookSparqlEndpointhook will be exposed. This will be used for waiting until the endpoint is available.queryRunnerReplication: Number of replication runs forsparql-benchmark-runner.queryRunnerWarmupRounds: Number of warmup runs forsparql-benchmark-runner.queryRunnerRequestDelay: Delay between requests in milliseconds forsparql-benchmark-runner.queryRunnerEndpointAvailabilityCheckTimeout: Timeout in milliseconds for checking if the SPARQL endpoint is available.queryRunnerUrlParams: A JSON record of string mappings containing URL parameters that will be passed to the SPARQL endpoint.queryTimeoutFallback: An optional timeout value for a single query in milliseconds, to be used as fallback in case the SPARQL endpoint hook's timeout fails. This should always be higher than the timeout value configured in the SPARQL endpoint hook.
License
jbr.js is written by Ruben Taelman.
This code is copyrighted by Ghent University – imec and released under the MIT license.
