Schema-Grounded NL→KQL: An Intermediate Representation for SIEM Rule Generation

An explicit, schema-validated intermediate representation between natural language and KQL that cuts field hallucination in LLM-generated Microsoft Sentinel detection rules from 93% to 13%.

SOC analysts routinely translate natural-language detection requirements — a line in an SOP, a sentence in a threat-intel report, an analyst's shorthand — into executable Kusto Query Language (KQL) rules for Microsoft Sentinel. Direct LLM generation of KQL is unreliable: it produces invalid syntax, hallucinated field names, and logically incorrect filters at rates too high for unsupervised SOC use. This project tests a specific hypothesis — that inserting an explicit, schema-validated Intermediate Representation (IR) between the natural language and the KQL reduces hallucination compared to direct generation — and measures how much, where, and why, on a purpose-built dataset derived from Microsoft's public Azure/Azure-Sentinel repository.

The idea is borrowed from compiler design: source code isn't translated straight to machine code — it passes through an IR, a structured, language-agnostic form that captures logic independently of source syntax and target instructions, and makes validation tractable. Instead of asking an LLM to go straight from 'detect credential stuffing' to raw KQL, the system first produces a typed, schema-checkable JSON object — the Security IR — describing the detection logic, validates it against the ASIM schema, and only then deterministically compiles it into KQL via templates. The bet: separate 'does the model understand the threat' from 'does the model remember KQL syntax', and solve the second with deterministic code rather than free-form generation. A newer capability scans each request for ambiguity and, rather than guess, asks a targeted clarifying question or abstains — so an underspecified prompt yields a question, not a confidently wrong rule.

Key metrics

  • Field Validity Rate — IR-mediated (System B): 86.7%
  • Field Validity Rate — direct baseline (System A): 6.7%
  • Syntax Validity / completion (System B): 97.8%
  • Repair Recovery Rate (≤3 attempts): 96.2%
  • Held-out completion (median, N=5): 88.9%
  • Held-out Logic Correctness (median): 82.4% (IQR 5.9)

Engineering rigor

  • n=45 primary A/B eval
  • McNemar paired significance
  • N=5 held-out replications
  • κ=0.645 inter-rater agreement

Tech stack

Python 3.11+, LangGraph, Pydantic (typed IR), Jinja2 (KQL templates), TF-IDF retrieval (RAG), pandas (execution substitute), Microsoft Sentinel / KQL, ASIM schema, McNemar's test / bootstrap CIs

Roadmap

  • A human (not just AI) rater pass on Logic Correctness (the scorer is built, not yet run).
  • Multi-platform templates (Sigma, SPL) consuming the same vendor-neutral IR.
  • Real-engine telemetry execution validation against a Kusto emulator or seeded workspace.
  • Optional MITRE ATT&CK technique/tactic tagging on generated rules.

View on GitHub

← All projects