Auto Parser
The Auto Parser uses a brute-force strategy: it iterates through every log-type record in the internal dataset and chooses the regex and templates that best matches the provided logs.
Compared with Template Matcher approaches, its key benefit is that you don’t need to supply templates or regex formatting during initialization, which makes it more convenient for rapid deployments. Its main drawback is that it only performs well for log types that are already included in the internal dataset.
The built-in dataset of log types cannot be modified by users and currently supports: HDFS, BGL, Audit, Syslog, OpenVPN, DNSmasq, and Apache.
It wraps functionality from the DetectMatePerformance project.
In/out
Input and output schemas in the pipeline
| Schema | Description | |
|---|---|---|
| Input | LogSchema | Unstructured log |
| Output | ParserSchema | Structured log |
Examples
Without fixing log type:
import yaml
from detectmatelibrary.parsers.autoparser import AutoParser
from detectmatelibrary.helper.from_to import From
with open("docs/examples/parsers/auto_parser.yaml") as f:
config = yaml.safe_load(f)
parser = AutoParser(name="AutoParser", config=config)
for j, parsed_log in enumerate(From.log(parser, "tests/test_data/audit.log")):
if j == 15:
break
print(parsed_log["template"]) # pid <*> uid <*> auid <*> ses <*> msg op <*> acct <*> exe <*> ...
With fixing log type:
# the same configuration, but skip the detection and fix the log type to Audit
config["parsers"]["AutoParser"]["params"]["fix_type"] = "Audit"
parser = AutoParser(name="AutoParser", config=config)
for j, parsed_log in enumerate(From.log(parser, "tests/test_data/audit.log")):
if j == 15:
break
print(parsed_log["template"]) # same template, without the type detection
Configuration file
The configuration used by the examples above. It sets only what this use case needs; every other parameter keeps its default (see Configuration arguments).
parsers:
AutoParser:
method_type: auto_parser
params:
data_use_training: 10 # the first 10 logs are used to recognise the log type
The same file works unchanged in both places a parser runs:
- Library: load it with
yaml.safe_loadand pass the dict asconfig=, as in the example. The key underparsers:must match the parser'sname. - DetectMateService: use it as the service's parser configuration.
Configuration arguments
All parameters this parser accepts, grouped by the YAML block they go in. Scope tells whether a parameter is specific to this parser or shared with other parsers (see the Parsers overview).
Top level
| Field | Type | Default | Scope | Description |
|---|---|---|---|---|
method_type |
string | auto_parser | shared | fitting description yet to find |
auto_config |
boolean | False | shared | Runs the configuration step before the training process. |
params
| Field | Type | Default | Scope | Description |
|---|---|---|---|---|
fix_type |
string | specific | fitting description yet to find | |
start_id |
integer | 10 | shared | Number used to start the unique ID generator. |
data_use_training |
integer, null | None | shared | Data used for training, if None, training is not done. |
data_use_configure |
integer, null | None | shared | Data used for configuration, if None, configuration is not done. |
use_config_data_as_training |
boolean | True | shared | Combine the configured data in the training process if True. |
train_buffer_max_records |
integer | 100000 | shared | Configure records kept in memory for training (use_config_data_as_training) before the buffer spills to Parquet files on disk, in parts of this many records. |
train_buffer_dir |
string, null | None | shared | Local directory for the spilled training buffer. None uses the system temp directory (TMPDIR). Each spill goes to a private detectmate-train-* directory, removed after training reads it; a killed process leaves it behind. |
log_format |
string, null | None | shared | fitting description yet to find |
time_format |
string, null | None | shared | fitting description yet to find |