First Draft: Nominal Sum Types via `enum union` and `switch` Expressions

Meta jared771 at gmail.com
Sun Sep 20 02:10:15 UTC 2026


On Friday, 18 September 2026 at 12:46:32 UTC, Dukc wrote:
> On Monday, 14 September 2026 at 06:35:13 UTC, Meta wrote:
>> This DIP proposes 2 new constructs for the D language: `enum 
>> union` as a language-level discriminated union type, and 
>> switch expressions which are used to inspect these unions at 
>> runtime.
>>
>> An enum union is declared using the `enum union` keyword:
>> ```d
>> enum union NetworkPacket
>> {
>>     case Data(const(ubyte)[]),
>>     case Ping(ulong),
>>     case Reset(ushort),
>>     case Heartbeat(),
>>     case EndOfStream(),
>> }
>> ```
>>
>> Variant declarations in an enum union must start with the 
>> keyword `case`; they represent one of the possible values that 
>> an enum union may take on. There are different types of 
>> variants that serve different functions.
>
> Like Nick, I think a semicolon separator is a better idea. This 
> will let `static if` to be easily used. It's especially 
> important for sum types, as a template parameter combined with 
> `static if`s in the body would be the D way to do [generalized 
> algebraic 
> datatypes](https://en.wikipedia.org/wiki/Generalized_algebraic_data_type).
>
> Having the possible variants declared with the same syntax as 
> union fields are is not a bad option. Your `case` prefix still 
> has its benefits though: it will make it easy to separate 
> actual variants from member functions, alternative constructors 
> and other fields.
>
> Style note: I think the tuple variants should be named 
> lower-case in teh standard D style. There's an argument for 
> naming them with upper initial because you might want to 
> distinguish pattern-matchable values from other and that's how 
> Haskell and OCaml work. But, the same could be said for 
> enumerated type values, which we already name with lower 
> initials. Therefore, lower initials for sum type variants too.
>
>
>> ```d
>> enum union DrawCommand
>> {
>>     // Unnamed positional parameters
>>     case MoveTo(double, double),
>>     case LineTo(double, double),
>>
>>     // Named positional parameters
>>     case Circle(double x, double y, double radius),
>>     case Text(string content, double x, double y, ubyte 
>> fontSize),
>>
>>     // Mixed named and unnamed positional parameters
>>     case Arc(double, double, double radius, double sweepAngle),
>> }
>> ```
>
> Additional idea: maybe you could optionally specify the tag 
> value like this:
>
> ```d
> enum union DrawCommand
> {
>     // Unnamed positional parameters
>     case MoveTo(double, double): 5;
>     case LineTo(double, double): 10;
>
>     // Named positional parameters
>     case Circle(double x, double y, double radius): 20;
>     case Text(string content, double x, double y, ubyte
> fontSize): 40;
>
>     // Mixed named and unnamed positional parameters
>     case Arc(double, double, double radius, double sweepAngle);
> }
> ```
>
> This could also let the tag field to be something else than an 
> ubyte. But I agree the default should be `ubyte`, and failing 
> that an `ushort` if nothing else is explicitly specified.

I thought about doing this with a compiler recognized @tag UDA, 
but I'm not convinced it'll see much use. It could easily be 
introduced in a subsequent DIP.

>> Struct variants are declared as a normal struct declaration:
>> ```d
>> enum union PaymentEvent
>> {
>>     case CardCharge {
>>         string token;
>>         ulong amountCents;
>>         string currency;
>>         bool require3DSecure;
>>     },
>>
>>     case BankTransfer {
>>         string iban;
>>         string bic;
>>         ulong amountCents;
>>         string reference;
>>     },
>>
>>     case RefundIssued {
>>         ulong originalTxId;
>>         ulong refundAmountCents;
>>         string reason;
>>     },
>> }
>> ```
>>
>> Note that struct variants are only allowed to declare fields; 
>> not methods, constructors, destructors, or any other type of 
>> declaration.
>>
>> Unlike tuple variants, struct variants are type declarations. 
>> They're also subtypes of the enum union:
>> ```d
>> auto charge = PaymentEvent.CardCharge("my token", 
>> 1_000_000_000, "CAD", true);
>> assert(is(typeof(charge): PaymentEvent));
>> ```
>
> I'm a bit worried this would be like `[1, 2, 3]` being typed as 
> `int[3]`, instead of `int[]` it is, for mostly good reasons.
>
> On the other hand, if `PaymentEvent.CardCharge("my token", 
> 1_000_000_000, "CAD", true)` would be `PaymentEvent` by default 
> that would be inconsistent too, since no other constructor call 
> returns a result of another type.
>
> Maybe we can initialise sum types with struct variants like 
> `PaymentEvent(CardCharge = {"my token", 1_000_000_000, "CAD", 
> true})` for the `PaymentEvent` type.

You can always use tuple variants instead, which are *not* their 
own type.

>> Bare-type variants directly embed an external type as a case 
>> in the enum union without needing to wrap it in a tuple 
>> variant:
>> ```d
>> enum union ConfigValue
>> {
>>     bool,
>>     long,
>>     double,
>>     string,
>>     string[],
>> }
>> ```
>
> I suggest these should still reguire the `case` prefix, 
> assuming other variants are declared with those too but you 
> decide to adapt the semicolon ending.

Ya this is from an older version of the DIP that accidentally got 
left in.

> An idea: Maybe `case bool;` could alternatively be written as 
> `case this(bool);`. Then `this(int){}` would be a regular 
> constructor function you'll define yourself and which will 
> probably call a `case` constructor.

Enum unions can already have constructors, so you can define 
`this(bool)` and `this(int)` if you want.

>> When it is ambiguous which type would be initialized by this 
>> assignment, the compiler requires the user to disambiguate:
>> ```d
>> enum union Nums
>> {
>>     case int,
>>     case long,
>> }
>>
>> Nums n = 0; // Error, 0 is ambiguous between variants `int` 
>> and `long` of enum union `Nums`
>> Nums n = 0L; // Ok
>> ```
>
> No, I think this should work the same way as function overload 
> resolution. `Nums n = 0;` should compile, since it's an exact 
> match to one of the variants (`case int`). If the fields were 
> `case long` and `case ulong`, then the first assignment would 
> be an ambiguity error but the second one would still compile.

Yeah I agree. This is a bad example. A better one is
```d
enum union Pointers
{
     case int*;
     case void*;
}

Pointers p = null; // Error, null is ambiguous between variants 
'int*' and 'void*'
```

> By the way, why would this overloading be limited to bare 
> variants? Shouldn't we just as well be able to write
> ```d
> enum union Integer
> {
>     case numeric(int),
>     case numeric(long),
>     case alphabetic(char),
>     case alphabetic(wchar),
>     case alphabetic(dchar)
> }
> ```
> ?

This doesn't work for a few reasons:

1. A variant name is not just a factory function—it is a runtime 
discriminant identifier. If numeric represents two different 
payload layouts and two different integer tags under the hood, 
the name ceases to identify the variant.

2. With your suggestion, this code now becomes ambiguous:
```d
Integer val = ...;
switch (val)
{
     case numeric(x) => ... // What is typeof(x)? int or long?
}
```

However, if you want functionality like this, you can do:
```d
enum union AnyOf(T...)
{
     static foreach (V; T)
         case T;
}

enum union Integer
{
     case numeric(AnyOf!(int, long));
     case alphabetic(AnyOf!(char, wchar, dchar));
}
```

And now here's the cool part: via the implicit construction rule, 
this works seamlessly:
```d
auto n = Integer.alphabetic('c'w); // Implicitly constructs 
AnyOf!(char, wchar, dchar) with char variant
```

>>
>> ## Enum Union Members
>>
>> Enum unions are treated as struct declarations internally, 
>> which contain a union with the declared variant cases, and a 
>> `__tag` value to track which variant is currently active.
>
> I have a different take on this. The type is called an `union`. 
> It's thus confusing if it behaves like a struct.
>
> Instead, I propose that sumtypes will be actual union 
> declarations. No `enum` needed before the `union`. For 
> constructors and fields prefixed with `case`, space will be 
> reserved for the tag along with the field. Other fields will 
> overlap the whole bit space, including the tag. For any union 
> with `case` fields, there will be an implicit untagged field 
> `ubyte casetag;`
>
> How do we ensure safety if those field types are mixed? My 
> suggestion: what happens depends on the default value, e.g. 
> whether a `case` or non-`case` field declaration comes first. 
> If it's `case`, all non-`case` fields are unsafe to modify, so 
> safe code can always rely on the tag staying valid for this 
> type. In the reverse case, all `case` fields with potentially 
> unsafe values are `@system`. Also, only if a `case` field comes 
> first, the union can have a destructor too like you propose 
> later.

This is outside my vision for the feature, and would be a 
completely separate DIP.

>>
>> Every enum union provides a `.init` value, which by default is 
>> the `.init` value of its first declared variant (in syntactic 
>> order). If that variant has an `@disable`'d init value, then 
>> the `.init` value of the second variant will be used. If all 
>> variants disable `.init`, the enum union will also have a 
>> disabled `.init`.
>
> My testing seems to show you can't actually `@disable` (or 
> override) an `.init` value in the sense that a parent type 
> would use it for it's own `.init` value. I'm not sure if that's 
> intentional, the spec doesn't say either way.

Ya I need to take that part out. I thought you could disable it 
for some reason.

>> ```d
>> enum union Option(T)
>> {
>>     case None,
>>     case Some(T),
>> }
>>
>> Option!int opt;
>> assert(opt.__tag == 0);
>> assert(opt == Option!int.None);
>> ```
>>
>> * If all types in the enum union disable default construction 
>> (`@disable this();`), default construction will be disabled 
>> for the union as well.
>
> The default constructor should be disabled straight away if one 
> for the first field is. This is how unions behave.

That's how it was initially, but then I realized that if one type 
in the union disables default construction, you can just use the 
next one in the list. So it's only necessary if _all_ of them 
disable it, unless I'm missing something.

>>
>> The enum union guarantees memory integrity across variant 
>> transformations:
>
> Good thinking here, I think your formulation works.
>
>
>>
>> * **RAII Lifecycle Dispatch**: If any variant contains an 
>> elaborate destructor (`~this()`), the compiler synthesizes an 
>> aggregate destructor that inspects the discriminant tag and 
>> invokes the destructor of the active variant.
>
> What happens if the tag is made invalid by `@system` 
> shenanigans? I suggest nothing, because destructors can't be 
> manually skipped.

If you're doing @system stuff, all bets are off and it's on you 
to maintain language invariants.

>>
>> When an enum union contains unit variants alongside 
>> non-nullable references, pointers (`T*`), class references, or 
>> bounded scalars (such as `bool`), the compiler exploits 
>> invalid bit patterns to encode the unit state:
>>
>> * `Option!(int*)`: The null pointer address `0x0` represents 
>> `None`.
>> * `sizeof(Option!(int*)) == 8` (on 64-bit platforms), 
>> incurring zero byte overhead for the tag.
>
> I'd leave this out. Predictable size and layout are more 
> valuable IMO.

Rust and Swift do this too, and it's a useful optimization in a 
lot of cases, so I think it should stay in.

>> ```d
>> string describePacket(NetworkPacket pkt)
>> {
>>     return switch (pkt)
>>     {
>>         // Positional payload extraction binding variables by 
>> value
>>         case Data(bytes)     => format("Data payload (%d 
>> bytes)", bytes.length),
>>         case Ping(timestamp) => format("Ping probe: 
>> timestamp=%d", timestamp),
>>         case Reset(code)     => format("Connection reset with 
>> code %d", code),
>>
>>         // Parameterless unit variants match directly by tag 
>> name
>>         case Heartbeat       => "Keep-alive heartbeat 
>> received",
>>         case EndOfStream     => "End of transmission stream",
>>     };
>> }
>> ```
>
> What if the case parameter is a manifest constant declared 
> somewhere and the user means to construct a `NetworkPacket` to 
> use as the switch case? One possibility: this would be done by 
> adding empty parenthesis, e.g. `case Data(magicValue)() => 
> "Magic data payload"`.

I'm not sure exactly what you mean, but symbols introduced by 
destructuring will shadow outer ones.

> Also, empty parenthesis should be required before `=>` for 
> parameterless variants. Otherwise `HeartBeat` and `EndOfStream` 
> are indistinquishable from untyped variables at the parser 
> level. Read ahead for reasons why we might want to have cases 
> with untyped variables.
>
>> Variable patterns are of the form `case Type name =>`. Their 
>> syntax mirrors the declaration of a local variable. These 
>> patterns can be used for any type of variant:
>> ```d
>> enum union A
>> {
>>     case int,
>>     case Unit(),
>>     case Tuple(int n, string),
>>     case Struct { double d; bool b; },
>>     case MyStruct = ExternalStruct,
>> }
>>
>> switch (A.Tuple(42, "asdf"))
>> {
>>     case int n => ...,
>>     case Unit u => ...,
>>     case Tuple t => ...,
>>     case Struct s => ...,
>>     case MyStruct m => ...,
>> }
>> ```
>
> In function literal declarations, `x => y` is the same as `(x) 
> => y`. This syntax doesn't follow that analogy, making things 
> confusing:  While `case int (n) => ...` still makes sense, 
> `case Unit (u) => ...` and `case Tuple (t) => ...` would 
> mismatch on number of tuple parameters. You want `case (Unit u) 
> => ...` and `case (Tuple t)` instead.
>
> We will probably also want to have "case-cases", and wildcards
>
> ```d
> enum union A
> {
>     case int, case Option!float, case string, case Exception;
>
>     static A someInstance;
> }
>
> switch(A.someInstance)
> {
>     case 42 () => ...,
>     // matches ints other than 42
>     case (int x) => ...,

This can be done with guard patterns:
case int n if (n == 42) => ...,
case int n => ..., // Catch-all int case

>     case Option!float None() => ...,
>     case Option!float Some x => ...,
>     // Wildcard, matches both strings and Exceptions, type 
> checks
>     // done for both.
>     case x => ...,
>     // Since a wildcard matcher preceeds, this is chosen only if
>     // the tag is invalid. Not required by the exhaustiveness
>     // checker.

The whole point of this feature is that the tag *can't* ever be 
invalid, modulo any @system stuff you do that violate language 
invariants.

>     default => ...
> }
> ```
> . Now, this is a big DIP already so it might be better to leave 
> this as a future possibility. But I think you should still have 
> rough plans how something like this could be done on top of 
> your DIP, while still keeping the syntax consistent.
>
>> Type Name patterns are the simplest form of pattern. They are 
>> of the form `case Type =>`, with no identifier. They are also 
>> supported for any type of variant:
>> ```d
>> switch (...)
>> {
>>     case int => ...,
>>     case Unit => ...,
>>     case Tuple => ...,
>>     case Struct => ...,
>>     case MyStruct => ...,
>> }
>> ```
>
> Assuming we're using function literals as the analogy, I'm 
> afraid this won't work. You can't write lambdas like this 
> either, because if you write `QuestionStruct => 42` the parser 
> will think that `QuestionStruct` is an untyped variable name, 
> not a type for an unnamed variable.

It does work and it parses just fine - I've already implemented 
it in the branch I link to in my original post.

>> Note that default arms may not access the active variant. The 
>> following is invalid:
>> ```d
>> switch (...)
>> {
>>     case int => ...,
>>     // default val => ..., Error
>> }
>> ```
>
> I'm not sure, maybe this should be supported. You would want to 
> do this if you're dealing with an invalid tag. If you buy my 
> union suggestion from earlier, we can even go a bit further:
> ```d
> switch (...)
> {
>     case int => ...,
>     default val => ...,
>     // alternatively
>     default casetag tag =>
> }
> ```
> E.g. The variant name after `default` refers to untagged union 
> field the variable will refer to. Without variant name, the 
> variable will refer to the union itself.
>
>> ```d
>> // Ok
>> cast(void)switch (pkt)
>> {
>>     case Data(bytes)     => format("Data payload (%d bytes)", 
>> bytes.length),
>>     case Ping(timestamp) => format("Ping probe: timestamp=%d", 
>> timestamp),
>>     case Reset(code)     => format("Connection reset with code 
>> %d", code),
>>     case Heartbeat       => "Keep-alive heartbeat received",
>>     case EndOfStream     => "End of transmission stream",
>> }; // Ending semicolon required
>> ```
>
> Regarding semicolons, I think they would be better case 
> delimiters between cases than commas, for the same reasons I 
> suggested that for the type declaration. Unlike a statement 
> switch, the scope of an expression switch would be a 
> declaration scope, not a function scope, so you couldn't have 
> control flow statements in it that wouldn't make sense anyway.

I don't want to do this because it then looks like the case arms 
are statements, when they're not - they're expressions.

> ## Overall thoughts
>
> Well, I'm sure suggesting a lot of changes, some pretty 
> ambitious. Maybe I got a bit carried away when writing this. My 
> suggestions probably have just as much holes as the existing 
> proposals after all.
>
> I'm convinced though, that your dip is an excellent starting 
> point. While I think it needs revision before it should be 
> adapted, it's because of the huge scope of the changes we're 
> talking about. Unlike most of the earlier DIPs on this I've 
> read, you're not speccing just a minimum vieable product but a 
> complete package. Despite that, it is closer to what I want 
> than many of those earlier DIPs. In other words, great work!

Thank you! 👌




More information about the dip.development mailing list