Skip to content

Auto Parser

The Auto Parser uses a brute-force strategy: it iterates through every log-type record in the internal dataset and chooses the regex and templates that best matches the provided logs.

Compared with Template Matcher approaches, its key benefit is that you don’t need to supply templates or regex formatting during initialization, which makes it more convenient for rapid deployments. Its main drawback is that it only performs well for log types that are already included in the internal dataset.

The built-in dataset of log types cannot be modified by users and currently supports: HDFS, BGL, Audit, Syslog, OpenVPN, DNSmasq, and Apache.

It wraps functionality from the DetectMatePerformance project.

In/out

Input and output schemas in the pipeline

Schema Description
Input LogSchema Unstructured log
Output ParserSchema Structured log

Examples

Without fixing log type:

import yaml
from detectmatelibrary.parsers.autoparser import AutoParser
from detectmatelibrary.helper.from_to import From

with open("docs/examples/parsers/auto_parser.yaml") as f:
    config = yaml.safe_load(f)
parser = AutoParser(name="AutoParser", config=config)

for j, parsed_log in enumerate(From.log(parser, "tests/test_data/audit.log")):
    if j == 15:
        break

print(parsed_log["template"])  # pid <*> uid <*> auid <*> ses <*> msg op <*> acct <*> exe <*> ...

With fixing log type:

# the same configuration, but skip the detection and fix the log type to Audit
config["parsers"]["AutoParser"]["params"]["fix_type"] = "Audit"
parser = AutoParser(name="AutoParser", config=config)

for j, parsed_log in enumerate(From.log(parser, "tests/test_data/audit.log")):
    if j == 15:
        break

print(parsed_log["template"])  # same template, without the type detection

Configuration file

The configuration used by the examples above. It sets only what this use case needs; every other parameter keeps its default (see Configuration arguments).

parsers:
  AutoParser:
    method_type: auto_parser
    params:
      data_use_training: 10   # the first 10 logs are used to recognise the log type

The same file works unchanged in both places a parser runs:

  • Library: load it with yaml.safe_load and pass the dict as config=, as in the example. The key under parsers: must match the parser's name.
  • DetectMateService: use it as the service's parser configuration.

Configuration arguments

All parameters this parser accepts, grouped by the YAML block they go in. Scope tells whether a parameter is specific to this parser or shared with other parsers (see the Parsers overview).

Top level
Field Type Default Scope Description
method_type string auto_parser shared fitting description yet to find
auto_config boolean False shared Runs the configuration step before the training process.
params
Field Type Default Scope Description
fix_type string specific fitting description yet to find
start_id integer 10 shared Number used to start the unique ID generator.
data_use_training integer, null None shared Data used for training, if None, training is not done.
data_use_configure integer, null None shared Data used for configuration, if None, configuration is not done.
use_config_data_as_training boolean True shared Combine the configured data in the training process if True.
train_buffer_max_records integer 100000 shared Configure records kept in memory for training (use_config_data_as_training) before the buffer spills to Parquet files on disk, in parts of this many records.
train_buffer_dir string, null None shared Local directory for the spilled training buffer. None uses the system temp directory (TMPDIR). Each spill goes to a private detectmate-train-* directory, removed after training reads it; a killed process leaves it behind.
log_format string, null None shared fitting description yet to find
time_format string, null None shared fitting description yet to find