First Draft: Nominal Sum Types via `enum union` and `switch` Expressions

Dukc ajieskola at gmail.com
Fri Sep 18 12:46:32 UTC 2026


On Monday, 14 September 2026 at 06:35:13 UTC, Meta wrote:
> This DIP proposes 2 new constructs for the D language: `enum 
> union` as a language-level discriminated union type, and switch 
> expressions which are used to inspect these unions at runtime.
>
> An enum union is declared using the `enum union` keyword:
> ```d
> enum union NetworkPacket
> {
>     case Data(const(ubyte)[]),
>     case Ping(ulong),
>     case Reset(ushort),
>     case Heartbeat(),
>     case EndOfStream(),
> }
> ```
>
> Variant declarations in an enum union must start with the 
> keyword `case`; they represent one of the possible values that 
> an enum union may take on. There are different types of 
> variants that serve different functions.

Like Nick, I think a semicolon separator is a better idea. This 
will let `static if` to be easily used. It's especially important 
for sum types, as a template parameter combined with `static if`s 
in the body would be the D way to do [generalized algebraic 
datatypes](https://en.wikipedia.org/wiki/Generalized_algebraic_data_type).

Having the possible variants declared with the same syntax as 
union fields are is not a bad option. Your `case` prefix still 
has its benefits though: it will make it easy to separate actual 
variants from member functions, alternative constructors and 
other fields.

Style note: I think the tuple variants should be named lower-case 
in teh standard D style. There's an argument for naming them with 
upper initial because you might want to distinguish 
pattern-matchable values from other and that's how Haskell and 
OCaml work. But, the same could be said for enumerated type 
values, which we already name with lower initials. Therefore, 
lower initials for sum type variants too.


> ```d
> enum union DrawCommand
> {
>     // Unnamed positional parameters
>     case MoveTo(double, double),
>     case LineTo(double, double),
>
>     // Named positional parameters
>     case Circle(double x, double y, double radius),
>     case Text(string content, double x, double y, ubyte 
> fontSize),
>
>     // Mixed named and unnamed positional parameters
>     case Arc(double, double, double radius, double sweepAngle),
> }
> ```

Additional idea: maybe you could optionally specify the tag value 
like this:

```d
enum union DrawCommand
{
     // Unnamed positional parameters
     case MoveTo(double, double): 5;
     case LineTo(double, double): 10;

     // Named positional parameters
     case Circle(double x, double y, double radius): 20;
     case Text(string content, double x, double y, ubyte
fontSize): 40;

     // Mixed named and unnamed positional parameters
     case Arc(double, double, double radius, double sweepAngle);
}
```

This could also let the tag field to be something else than an 
ubyte. But I agree the default should be `ubyte`, and failing 
that an `ushort` if nothing else is explicitly specified.

> Struct variants are declared as a normal struct declaration:
> ```d
> enum union PaymentEvent
> {
>     case CardCharge {
>         string token;
>         ulong amountCents;
>         string currency;
>         bool require3DSecure;
>     },
>
>     case BankTransfer {
>         string iban;
>         string bic;
>         ulong amountCents;
>         string reference;
>     },
>
>     case RefundIssued {
>         ulong originalTxId;
>         ulong refundAmountCents;
>         string reason;
>     },
> }
> ```
>
> Note that struct variants are only allowed to declare fields; 
> not methods, constructors, destructors, or any other type of 
> declaration.
>
> Unlike tuple variants, struct variants are type declarations. 
> They're also subtypes of the enum union:
> ```d
> auto charge = PaymentEvent.CardCharge("my token", 
> 1_000_000_000, "CAD", true);
> assert(is(typeof(charge): PaymentEvent));
> ```

I'm a bit worried this would be like `[1, 2, 3]` being typed as 
`int[3]`, instead of `int[]` it is, for mostly good reasons.

On the other hand, if `PaymentEvent.CardCharge("my token", 
1_000_000_000, "CAD", true)` would be `PaymentEvent` by default 
that would be inconsistent too, since no other constructor call 
returns a result of another type.

Maybe we can initialise sum types with struct variants like 
`PaymentEvent(CardCharge = {"my token", 1_000_000_000, "CAD", 
true})` for the `PaymentEvent` type.

> Bare-type variants directly embed an external type as a case in 
> the enum union without needing to wrap it in a tuple variant:
> ```d
> enum union ConfigValue
> {
>     bool,
>     long,
>     double,
>     string,
>     string[],
> }
> ```

I suggest these should still reguire the `case` prefix, assuming 
other variants are declared with those too but you decide to 
adapt the semicolon ending.

An idea: Maybe `case bool;` could alternatively be written as 
`case this(bool);`. Then `this(int){}` would be a regular 
constructor function you'll define yourself and which will 
probably call a `case` constructor.

> When it is ambiguous which type would be initialized by this 
> assignment, the compiler requires the user to disambiguate:
> ```d
> enum union Nums
> {
>     case int,
>     case long,
> }
>
> Nums n = 0; // Error, 0 is ambiguous between variants `int` and 
> `long` of enum union `Nums`
> Nums n = 0L; // Ok
> ```

No, I think this should work the same way as function overload 
resolution. `Nums n = 0;` should compile, since it's an exact 
match to one of the variants (`case int`). If the fields were 
`case long` and `case ulong`, then the first assignment would be 
an ambiguity error but the second one would still compile.

By the way, why would this overloading be limited to bare 
variants? Shouldn't we just as well be able to write
```d
enum union Integer
{
     case numeric(int),
     case numeric(long),
     case alphabetic(char),
     case alphabetic(wchar),
     case alphabetic(dchar)
}
```
?

>
> ## Enum Union Members
>
> Enum unions are treated as struct declarations internally, 
> which contain a union with the declared variant cases, and a 
> `__tag` value to track which variant is currently active.

I have a different take on this. The type is called an `union`. 
It's thus confusing if it behaves like a struct.

Instead, I propose that sumtypes will be actual union 
declarations. No `enum` needed before the `union`. For 
constructors and fields prefixed with `case`, space will be 
reserved for the tag along with the field. Other fields will 
overlap the whole bit space, including the tag. For any union 
with `case` fields, there will be an implicit untagged field 
`ubyte casetag;`

How do we ensure safety if those field types are mixed? My 
suggestion: what happens depends on the default value, e.g. 
whether a `case` or non-`case` field declaration comes first. If 
it's `case`, all non-`case` fields are unsafe to modify, so safe 
code can always rely on the tag staying valid for this type. In 
the reverse case, all `case` fields with potentially unsafe 
values are `@system`. Also, only if a `case` field comes first, 
the union can have a destructor too like you propose later.

>
> Every enum union provides a `.init` value, which by default is 
> the `.init` value of its first declared variant (in syntactic 
> order). If that variant has an `@disable`'d init value, then 
> the `.init` value of the second variant will be used. If all 
> variants disable `.init`, the enum union will also have a 
> disabled `.init`.

My testing seems to show you can't actually `@disable` (or 
override) an `.init` value in the sense that a parent type would 
use it for it's own `.init` value. I'm not sure if that's 
intentional, the spec doesn't say either way.

> ```d
> enum union Option(T)
> {
>     case None,
>     case Some(T),
> }
>
> Option!int opt;
> assert(opt.__tag == 0);
> assert(opt == Option!int.None);
> ```
>
> * If all types in the enum union disable default construction 
> (`@disable this();`), default construction will be disabled for 
> the union as well.

The default constructor should be disabled straight away if one 
for the first field is. This is how unions behave.

>
> The enum union guarantees memory integrity across variant 
> transformations:

Good thinking here, I think your formulation works.


>
> * **RAII Lifecycle Dispatch**: If any variant contains an 
> elaborate destructor (`~this()`), the compiler synthesizes an 
> aggregate destructor that inspects the discriminant tag and 
> invokes the destructor of the active variant.

What happens if the tag is made invalid by `@system` shenanigans? 
I suggest nothing, because destructors can't be manually skipped.

>
> When an enum union contains unit variants alongside 
> non-nullable references, pointers (`T*`), class references, or 
> bounded scalars (such as `bool`), the compiler exploits invalid 
> bit patterns to encode the unit state:
>
> * `Option!(int*)`: The null pointer address `0x0` represents 
> `None`.
> * `sizeof(Option!(int*)) == 8` (on 64-bit platforms), incurring 
> zero byte overhead for the tag.

I'd leave this out. Predictable size and layout are more valuable 
IMO.

> ```d
> string describePacket(NetworkPacket pkt)
> {
>     return switch (pkt)
>     {
>         // Positional payload extraction binding variables by 
> value
>         case Data(bytes)     => format("Data payload (%d 
> bytes)", bytes.length),
>         case Ping(timestamp) => format("Ping probe: 
> timestamp=%d", timestamp),
>         case Reset(code)     => format("Connection reset with 
> code %d", code),
>
>         // Parameterless unit variants match directly by tag 
> name
>         case Heartbeat       => "Keep-alive heartbeat received",
>         case EndOfStream     => "End of transmission stream",
>     };
> }
> ```

What if the case parameter is a manifest constant declared 
somewhere and the user means to construct a `NetworkPacket` to 
use as the switch case? One possibility: this would be done by 
adding empty parenthesis, e.g. `case Data(magicValue)() => "Magic 
data payload"`.

Also, empty parenthesis should be required before `=>` for 
parameterless variants. Otherwise `HeartBeat` and `EndOfStream` 
are indistinquishable from untyped variables at the parser level. 
Read ahead for reasons why we might want to have cases with 
untyped variables.

> Variable patterns are of the form `case Type name =>`. Their 
> syntax mirrors the declaration of a local variable. These 
> patterns can be used for any type of variant:
> ```d
> enum union A
> {
>     case int,
>     case Unit(),
>     case Tuple(int n, string),
>     case Struct { double d; bool b; },
>     case MyStruct = ExternalStruct,
> }
>
> switch (A.Tuple(42, "asdf"))
> {
>     case int n => ...,
>     case Unit u => ...,
>     case Tuple t => ...,
>     case Struct s => ...,
>     case MyStruct m => ...,
> }
> ```

In function literal declarations, `x => y` is the same as `(x) => 
y`. This syntax doesn't follow that analogy, making things 
confusing:  While `case int (n) => ...` still makes sense, `case 
Unit (u) => ...` and `case Tuple (t) => ...` would mismatch on 
number of tuple parameters. You want `case (Unit u) => ...` and 
`case (Tuple t)` instead.

We will probably also want to have "case-cases", and wildcards

```d
enum union A
{
     case int, case Option!float, case string, case Exception;

     static A someInstance;
}

switch(A.someInstance)
{
     case 42 () => ...,
     // matches ints other than 42
     case (int x) => ...,
     case Option!float None() => ...,
     case Option!float Some x => ...,
     // Wildcard, matches both strings and Exceptions, type checks
     // done for both.
     case x => ...,
     // Since a wildcard matcher preceeds, this is chosen only if
     // the tag is invalid. Not required by the exhaustiveness
     // checker.
     default => ...
}
```
. Now, this is a big DIP already so it might be better to leave 
this as a future possibility. But I think you should still have 
rough plans how something like this could be done on top of your 
DIP, while still keeping the syntax consistent.

> Type Name patterns are the simplest form of pattern. They are 
> of the form `case Type =>`, with no identifier. They are also 
> supported for any type of variant:
> ```d
> switch (...)
> {
>     case int => ...,
>     case Unit => ...,
>     case Tuple => ...,
>     case Struct => ...,
>     case MyStruct => ...,
> }
> ```

Assuming we're using function literals as the analogy, I'm afraid 
this won't work. You can't write lambdas like this either, 
because if you write `QuestionStruct => 42` the parser will think 
that `QuestionStruct` is an untyped variable name, not a type for 
an unnamed variable.

> Note that default arms may not access the active variant. The 
> following is invalid:
> ```d
> switch (...)
> {
>     case int => ...,
>     // default val => ..., Error
> }
> ```

I'm not sure, maybe this should be supported. You would want to 
do this if you're dealing with an invalid tag. If you buy my 
union suggestion from earlier, we can even go a bit further:
```d
switch (...)
{
     case int => ...,
     default val => ...,
     // alternatively
     default casetag tag =>
}
```
E.g. The variant name after `default` refers to untagged union 
field the variable will refer to. Without variant name, the 
variable will refer to the union itself.

> ```d
> // Ok
> cast(void)switch (pkt)
> {
>     case Data(bytes)     => format("Data payload (%d bytes)", 
> bytes.length),
>     case Ping(timestamp) => format("Ping probe: timestamp=%d", 
> timestamp),
>     case Reset(code)     => format("Connection reset with code 
> %d", code),
>     case Heartbeat       => "Keep-alive heartbeat received",
>     case EndOfStream     => "End of transmission stream",
> }; // Ending semicolon required
> ```

Regarding semicolons, I think they would be better case 
delimiters between cases than commas, for the same reasons I 
suggested that for the type declaration. Unlike a statement 
switch, the scope of an expression switch would be a declaration 
scope, not a function scope, so you couldn't have control flow 
statements in it that wouldn't make sense anyway.

## Overall thoughts

Well, I'm sure suggesting a lot of changes, some pretty 
ambitious. Maybe I got a bit carried away when writing this. My 
suggestions probably have just as much holes as the existing 
proposals after all.

I'm convinced though, that your dip is an excellent starting 
point. While I think it needs revision before it should be 
adapted, it's because of the huge scope of the changes we're 
talking about. Unlike most of the earlier DIPs on this I've read, 
you're not speccing just a minimum vieable product but a complete 
package. Despite that, it is closer to what I want than many of 
those earlier DIPs. In other words, great work!


More information about the dip.development mailing list