First Draft: Nominal Sum Types via `enum union` and `switch` Expressions
Meta
jared771 at gmail.com
Mon Sep 21 14:17:51 UTC 2026
On Sunday, 20 September 2026 at 21:20:29 UTC, Dukc wrote:
> On Sunday, 20 September 2026 at 02:10:15 UTC, Meta wrote:
>>> I'm a bit worried this would be like `[1, 2, 3]` being typed
>>> as `int[3]`, instead of `int[]` it is, for mostly good
>>> reasons.
>>>
>>> On the other hand, if `PaymentEvent.CardCharge("my token",
>>> 1_000_000_000, "CAD", true)` would be `PaymentEvent` by
>>> default that would be inconsistent too, since no other
>>> constructor call returns a result of another type.
>>>
>>> Maybe we can initialise sum types with struct variants like
>>> `PaymentEvent(CardCharge = {"my token", 1_000_000_000, "CAD",
>>> true})` for the `PaymentEvent` type.
>>
>> You can always use tuple variants instead, which are *not*
>> their own type.
>
> I wasn't precise enough, I suppose. I wasn't worried about
> struct variants having their own type per se, I was worried
> about the implications of having that type _inferred by
> default_. Although I'm not actually proposing any changes here
> since I didn't come up with better ideas.
>
>>> An idea: Maybe `case bool;` could alternatively be written as
>>> `case this(bool);`. Then `this(int){}` would be a regular
>>> constructor function you'll define yourself and which will
>>> probably call a `case` constructor.
>>
>> Enum unions can already have constructors, so you can define
>> `this(bool)` and `this(int)` if you want.
>
> Yes, but the idea is that since tuple variants are declared
> like functions, you could do the same with bare type variants
> using the `this` keyword. Stylistic thing only, but would be
> consistent IMO.
>
>
>> Yeah I agree. This is a bad example. A better one is
>> ```d
>> enum union Pointers
>> {
>> case int*;
>> case void*;
>> }
>>
>> Pointers p = null; // Error, null is ambiguous between
>> variants 'int*' and 'void*'
>> ```
>
> Sorry, that is still wrong. Spec 20.14.6:
>> If two or more functions have the same match level, then
>> partial ordering is used to disambiguate to find the best
>> match. Partial ordering finds the most specialized function.
>> If neither function is more specialized than the other, then
>> it is an ambiguity error. Partial ordering is determined for
>> functions f and g by taking the parameter types of f,
>> constructing a list of arguments by taking the default values
>> of those types, and attempting to match them against g. If it
>> succeeds, then g is at least as specialized as f.
>
> This means that `int*` is more specialised than `void*` and
> should therefore be picked.
Ya, I meant for that to be long*, not void*.
>> 2. With your suggestion, this code now becomes ambiguous:
>> ```d
>> Integer val = ...;
>> switch (val)
>> {
>> case numeric(x) => ... // What is typeof(x)? int or long?
>> }
>> ```
>
> For now, you'd have to specify a type for `x`, just as with
> bare type variants. If a future DIP will implement wildcard
> matching, `x` will be a wildcard that matches both, probably
> either like a templated function or a `case` inside `static if`.
Rikki included it in his DIP, but I don't really see the point to
allowing a wildcard case. The whole idea is that you are
intentionally handling each case of the enum union. Also what
happens if you have a wildcard case *and* a default arm? And
multiple unhandled cases? Which arm captures which cases? I don't
think the extra complexity is worth it.
>>
>> This is outside my vision for the feature, and would be a
>> completely separate DIP.
>>
>
> Well you are already specifying your sum types are structs in
> disguise, which is likewise out of scope for what is minimally
> required for sum types.
How else would they be represented? A struct that contains a
union has been the canonical way to do this for the past 50
years, and it's a very easy, simple lowering.
> If implemented it will dictate the basic idea how sum types
> work. This is a high level strategic design decision that
> deserves a lot of scrutiny.
>
> But obviously, we do want to pick something and you're the DIP
> author so of course it should be something you find convincing.
> My personal opinion is that if sum types are another aggregate
> types under the hood, and their declarations say `union` on the
> tin, they should be unions.
They *are* unions. Tagged unions.
> I guess changing the keyword would be an okay alternative for
> me.
`enum union` is just a placeholder name, but it also makes the
most sense to me. They have aspects of both enums, with the
enumerated cases, and unions, with the overlapping storage.
Simple.
>> That's how it was initially, but then I realized that if one
>> type in the union disables default construction, you can just
>> use the next one in the list. So it's only necessary if _all_
>> of them disable it, unless I'm missing something.
>>
>
> It's not necessary for the language to work, but it's necessary
> to avoid an inconsistency with rest of the language. Users are
> probably going to expect that untagged and tagged unions behave
> similarily on matters like this.
That's true, but the reason it works like that for unions is
because they're untagged. The rule is unnecessary for tagged
unions.
>> If you're doing @system stuff, all bets are off and it's on
>> you to maintain language invariants.
>
> The question is, what are_those invariants in this case? You
> can specify that a sum type must have a valid tag on
> destruction or else undefined or unspecified behaviour will
> follow. But I don't think that's a good idea, since it isn't
> possible to skip the destructor call of a local variable
Named Return Value Optimization.
> (except by things like `longjmp` or foreign language
> exceptions). Hence my suggestion to specify the destructor will
> do nothing if the tag is invalid.
You're wildly over-complicating this. If the user wants to mess
around and do @system stuff with an enum union, they're on their
own.
>>>
>>> What if the case parameter is a manifest constant declared
>>> somewhere and the user means to construct a `NetworkPacket`
>>> to use as the switch case? One possibility: this would be
>>> done by adding empty parenthesis, e.g. `case
>>> Data(magicValue)() => "Magic data payload"`.
>>
>> I'm not sure exactly what you mean, but symbols introduced by
>> destructuring will shadow outer ones.
>>
>
> I basically mean that if you have `switch (sum) { case x(y) =>
> smth, ...}`, it could either mean that `x` is a variant and `y`
> is a variable declaration for `smth`, or that `x` is a function
> or a constructor called with value `y`, and `sum` is matched
> against the result (as in existing switch statements).
It's invalid to have a function call here, so there's no
ambiguity. Even if this DIP introduced full pattern matching,
that wouldn't be valid.
> For the purposes of this DIP specifically, maybe these could be
> distinguished by the type of `sum` - always the first meaning
> if a sum type, always the second meaning if something else. But
> on this point, I think we need to look further than that. More
> on that on the next point.
>
>>
>> This can be done with guard patterns:
>> case int n if (n == 42) => ...,
>> case int n => ..., // Catch-all int case
>
> Yes. But many other languages with pattern matching do support
> what I propose here without needing guards. It's likely we will
> want to support the same in the future.
>
> Like I wrote earlier, I don't think your DIP necessarily needs
> to propose supporting that. But, a future DIP will have to
> build on the syntax we nail down now. We need to think already
> how the more advanced forms of pattern matching could be done
> on syntactic level, otherwise the future DIP will have to
> settle on ugly workarounds.
I've already put a significant amount of thought into it, and it
works seamlessly with the pattern syntax introduces in this DIP.
> This is why I insist on the rule that if an identifier comes
> before `=>` (or the guard pattern), it must always be a
> variable declaration (not a variant name, nor a type for bare
> type matching), and absent a variable declaration there must be
> an empty pair of parenthesis.
>
>> The whole point of this feature is that the tag *can't* ever
>> be invalid, modulo any @system stuff you do that violate
>> language invariants.
>
> I agree the tag shouldn't ever be invalid in `@safe` code (at
> least for some sum types), but D is a systems programming
> language. The language should support creative bit manipulation
> over the types, as long as the user properly protects the
> invariants of those types manually. For example, if the tag
> says the sum type is an array, the user has the responsibility
> to either store an actual slice in it, or to prevent use of the
> invalid data.
>
>>> Assuming we're using function literals as the analogy, I'm
>>> afraid this won't work. You can't write lambdas like this
>>> either, because if you write `QuestionStruct => 42` the
>>> parser will think that `QuestionStruct` is an untyped
>>> variable name, not a type for an unnamed variable.
>>
>> It does work and it parses just fine - I've already
>> implemented it in the branch I link to in my original post.
>>
>
> By doing this, you're closing off the possiblity - or at least
> making it indistinquishable for the parser - for `case x => 42`
> to mean a wildcard match in the future.
I'm fine with that. If you think it's worthwhile, you will need
to show me a very convincing example of why it's useful.
> Admittedly, the future DIP could use a different syntax for a
> wildcard match but this seems the best option for me.
More information about the dip.development
mailing list