How to update a DMARC aggregate-report parser for RFC 9990

dmarctutorial

An RFC 7489 DMARC aggregate-report parser can appear to work after the DMARC specification moves on. It still finds feedback, record, source_ip, and policy_evaluated. The danger is that it silently loses the meaning of newer reports: a namespace check fails, a policy-discovery field is ignored, or a database collapses distinct policies into one row.

RFC 9990, published as the aggregate-reporting companion to RFC 9989, obsoletes RFC 7489. The XML is still recognizably DMARC, but a production parser needs to update its namespace handling, policy model, validation rules, and deduplication assumptions.

This is a parser migration guide, not a recommendation to throw away every old report. Existing RFC 7489 reports remain useful historical input. The goal is to accept both generations deliberately, preserve what the receiver reported, and avoid pretending that an old report contains fields it never had.

What changed in the report model

The important changes are concentrated in policy_published and in the document namespace.

AreaOlder reportsRFC 9990 reports
XML namespaceCommonly urn:ietf:params:xml:ns:dmarc-1.0urn:ietf:params:xml:ns:dmarc-2.0
Policy discoveryUsually inferred from the report contextOptional discovery_method: psl or treewalk
Nonexistent-subdomain policyNot representedOptional np
Extension pointsLimited parser assumptions are commonNamespaced elements may appear at file and record level
Policy changes during a periodOften treated as impossibleMixed or separately reported configurations are explicitly possible

The fields are not cosmetic. A report saying discovery_method=treewalk tells the consumer that the receiver used the RFC 9989 DNS tree-walk model rather than the older public-suffix-list approach. The np value tells you what policy was available for a nonexistent subdomain, which is not interchangeable with sp.

Start with namespace-aware XML parsing

The most common migration bug is selecting elements by their unqualified names:

root.find("report_metadata")

That works only when the input happens not to use a default namespace and the XML library treats the element name as unqualified. An RFC 9990 document has a DMARC namespace on feedback, and that namespace applies to its unprefixed descendants.

Use the namespace URI as the identity of the vocabulary:

DMARC_NAMESPACES = {
    "urn:ietf:params:xml:ns:dmarc-1.0": "rfc7489",
    "urn:ietf:params:xml:ns:dmarc-2.0": "rfc9990",
}

def local_name(tag):
    return tag.rsplit("}", 1)[-1]

def namespace_uri(tag):
    return tag[1:].split("}", 1)[0] if tag.startswith("{") else None

namespace = namespace_uri(root.tag)
if namespace not in DMARC_NAMESPACES:
    raise InvalidReport("unsupported DMARC namespace")

def child(parent, name):
    return parent.find(f"{{{namespace}}}{name}")

The exact API differs between XML libraries, but the rules should stay the same:

  1. Parse the namespace URI from the root element.
  2. Accept the RFC 7489 and RFC 9990 namespaces if backward compatibility is intentional.
  3. Reject an unknown DMARC namespace, or quarantine it for manual review.
  4. Use a qualified name for standard DMARC children.
  5. Do not treat a different namespace as a DMARC field just because its local name is policy_published or record.

Do not “fix” this by stripping namespaces from every element before parsing. That makes extension elements indistinguishable from standard fields and can let a namespaced extension overwrite data in the core report model.

Validate the root and the required structure

RFC 9990 defines feedback as the root and requires the core first-level elements in order: version is optional, followed by report_metadata, policy_published, an optional extension, and one or more record elements.

The order matters for schema validation. A hand-written parser does not need to reject every harmless formatting difference, but it should enforce the semantic requirements:

  • root is feedback in the supported DMARC namespace
  • report_metadata exists exactly once
  • policy_published exists exactly once
  • at least one record exists
  • each record has row, identifiers, and auth_results
  • required scalar fields are present and parseable
  • an absent optional field is different from an empty required field

The report's version, when present, must be 1.0. Do not use it as the namespace selector. The version value remains 1.0 while the XML namespace identifies the report vocabulary.

If you use an XSD validator, obtain the schema from the RFC 9990 Appendix A or its authoritative distribution and still keep application-level checks. Schema validation tells you whether the XML fits the document shape. It does not tell you whether the report is a duplicate, whether its timestamps overlap a previously ingested report, or whether a sender's policy changed halfway through a reporting window.

Map the new policy fields explicitly

The policy_published element now has this relevant shape:

<policy_published>
  <domain>example.com</domain>
  <discovery_method>treewalk</discovery_method>
  <p>reject</p>
  <sp>quarantine</sp>
  <np>reject</np>
  <fo>0</fo>
  <adkim>r</adkim>
  <aspf>r</aspf>
  <testing>n</testing>
</policy_published>

Store these as separate fields, rather than putting the complete XML fragment in a legacy policy string:

XML elementParser behavior
domainRequired policy domain. Preserve the normalized domain and, if useful, the original text.
discovery_methodOptional enum: psl or treewalk. Preserve null when absent.
pRequired in the report schema. Store the reported value, not a value inferred from your own DNS lookup.
spOptional existing-subdomain policy. Do not fill it with p during ingestion.
npOptional nonexistent-subdomain policy. Do not treat it as an alias for sp.
foOptional failure-reporting option string. Preserve it even if your product does not process failure reports.
adkim, aspfOptional alignment modes. Apply documented defaults only in a derived view.
testingOptional value of the DMARC t tag. Preserve the receiver's reported value.

The distinction between raw and derived values is important. RFC 9990 says unspecified tags have default values, but an absent sp is different evidence from an explicit sp=none. Keep the raw presence bit or raw nullable value, then calculate effective policy separately when a report is displayed.

Similarly, discovery_method describes how the reporting receiver discovered the policy. It is not permission for the report consumer to rerun DNS and replace the reported policy with today's answer. DNS may have changed since the message was evaluated.

Treat np as a first-class policy

sp covers existing subdomains without their own applicable record. np covers nonexistent subdomains, where the DNS response indicates NXDOMAIN. RFC 9989 defines the fallback for an absent np: use sp when present, otherwise p.

That fallback can be represented without destroying the source data:

def effective_subdomain_policy(policy, author_domain_exists):
    if author_domain_exists:
        return policy.sp if policy.sp is not None else policy.p
    return policy.np if policy.np is not None else (
        policy.sp if policy.sp is not None else policy.p
    )

This function is for interpretation, not ingestion. A report processor often does not have enough historical DNS information to independently prove whether the Author Domain existed at evaluation time. The report's policy_published data is evidence of the receiver's discovered configuration; store it before applying any UI-level explanation.

Keep policy domain separate from message domains

One report has one DMARC Policy Domain, but it can contain multiple record elements for different header_from domains. For example, a policy at example.com can produce records for both example.com and news.example.com.

Do not use identifiers/header_from as the report's policy domain. Keep at least these values separate:

  • report policy domain from policy_published/domain
  • message Author Domain from identifiers/header_from
  • SPF domain from auth_results/spf/domain, when present
  • each DKIM signing domain from auth_results/dkim/domain
  • the report generator and report_id

This separation becomes more valuable with tree-walk discovery. The policy record can be found at one domain while the Author Domain in a record is a subdomain that inherited or otherwise used that policy.

Do not collapse extensions into core fields

RFC 9990 provides an optional file-level extension element and namespaced extension elements after auth_results inside a record. A parser that rejects every unknown child will become brittle as extensions appear. A parser that accepts every unknown child as if it were core DMARC data can corrupt the model.

A practical policy is:

  1. Validate and map core DMARC elements in the DMARC namespace.
  2. Preserve unknown namespaced extensions as raw XML or a separate extension table.
  3. Ignore an extension you do not understand after recording its namespace and local name.
  4. Reject an unqualified unknown element where the schema does not permit one.
  5. Never let an extension replace a core p, record, dkim, or spf value.

For example, this is an extension, not a new core field:

<extension xmlns:ext="urn:example:mail-extension">
  <ext:arc-override>never</ext:arc-override>
</extension>

The namespace URI is the boundary. Prefix names such as ext are chosen by the document and must not be treated as stable identifiers.

Update deduplication and policy-change handling

The RFC 9990 report ID is intended to be unique among reports sent to the same domain. Use it as a primary deduplication signal together with the reporting organization and policy domain. Keep the attachment checksum as a secondary safeguard for malformed or repeated deliveries.

Do not deduplicate only on the date range. Two reports can cover the same interval and still be separate reports from different generators or for different policy configurations.

Also stop assuming that one reporting period has one effective policy. RFC 9990 describes two possible outcomes when DNS policy changes during the period:

  • one report can contain dispositions based on a mix of old and new policy observations, while exposing one policy_published value
  • the receiver can send multiple reports for the period, one for each observed configuration

The safest data model preserves the report as received and flags overlapping intervals for analysis. It should not rewrite every historical row to match the current DNS record.

Handle old and new reports during the transition

A useful migration strategy is to make the parser vocabulary-aware while keeping the normalized business model stable:

  1. Detect the root namespace and record the detected report format.
  2. Parse RFC 7489 and RFC 9990 core fields through the same semantic mapping.
  3. Populate discovery_method and np only when supplied by the report.
  4. Preserve a nullable value for fields unavailable in older reports.
  5. Keep the original XML or a redacted diagnostic sample for parser regression tests.
  6. Add fixtures for default namespaces, prefixed namespaces, unknown extensions, empty envelope_from, and multiple DKIM results.
  7. Compare aggregate totals before and after the migration by report ID and policy domain.

Do not infer that an RFC 7489 report used psl merely because it lacks discovery_method. The field did not exist in that report vocabulary. You may display “not reported” or derive a separate compatibility label, but do not present an inference as receiver evidence.

The namespace and schema identify the report vocabulary; discovery_method identifies the receiver's policy-discovery method. They answer different questions and should be stored separately.

A compact migration checklist

Before deploying the parser update, verify that it can:

  • accept urn:ietf:params:xml:ns:dmarc-2.0
  • continue handling the older namespace if historical imports require it
  • qualify all core XML lookups by namespace
  • parse optional discovery_method values psl and treewalk
  • parse and retain np without confusing it with sp
  • preserve absent optional fields as absent
  • keep policy domain, Author Domain, SPF domain, and DKIM domains separate
  • tolerate supported namespaced extensions without flattening them into core data
  • validate required structure and version=1.0
  • deduplicate using report identity rather than only time range
  • flag overlapping reports and possible policy changes instead of overwriting history

For a broader explanation of what the new discovery method means operationally, see How DMARC tree-walk policy discovery changes Organizational Domain and public-suffix handling. For the older XML fields and report-reading workflow, DMARC reporting 101 remains a useful starting point.

Make the report explainable

An RFC 9990 parser upgrade is successful when it preserves context, not merely when it stops throwing XML errors. Store the namespace, report format, policy domain, discovery method, raw policy-field presence, and extension data alongside the familiar authentication results.

That gives operators an answer to the questions that matter during a transition: what did the receiver discover, where was the policy found, which policy values were actually reported, and whether two apparently different results came from different discovery methods? Once those distinctions survive ingestion, the rest of the reporting interface can explain the data instead of guessing.

Previous Post