Reducing undefined behavior in the C language

(lwn.net)

120 points | by signa11 11 hours ago ago

128 comments

  • chasil 9 hours ago ago

    "Some Honeywell machines, for example, had nine-bit bytes."

    OS 2200 has 36-bit words. It is still a supported platform.

    https://en.wikipedia.org/wiki/UNIVAC_1100/2200_series

    This platform was the first SMP UNIX implementation:

    "Any configuration supplied by Sperry, including multiprocessor ones, can run the UNIX system."

    https://www.nokia.com/bell-labs/about/dennis-m-ritchie/other...

    • strenholme 5 hours ago ago

      The last Unisys release, the Dorado from 2015 (ClearPath Dorado 8300 to be exact), used Xeon processors (which are 64-bit x86_64). But, for argument’s sake, let’s suppose there’s a Unisys 2200 out there with 36-bit registers and GCC/clang makes a C2y compiler for it.

      If so, then said compiler should still have support for (u)int_8/16/32/64_t. Said compiler already needs to have that support with (unsigned) _BitInt(8/16/32/64)—and, yes, one can also have 36-bit ints on x86_64 in C23 with _BitInt(36) if one must—and allowing stuff like int32_t and uint32_t will allow said (imaginary) 2200 system to cleanly compile a lot of pre-C23 open source code out there. Yes, uint32_t will look a little ugly at the assembly level, just as _BitInt(36) will look a little ugly on, say, a Xeon processor, but the code will compile and run the same.

      Of course, these days they can buy tokens so an agent can do any relevant porting, but still.

    • eru 9 hours ago ago

      > OS 2200 has 36-bit words. It is still a supported platform.

      Yes, but that's perhaps an argument for 'implementation defined behaviour', not in favour of 'undefined behaviour'.

  • pizlonator 7 hours ago ago

    It’s cool that this mentions Fil-C but it also undersells it. TFA also undersells CHERI. Fil-C doesn’t just “find a lot of temporal-safety” bugs. It closes off memory safety bugs (special and temporal) for exploit writers and ascribes a tight semantics to the whole language. CHERI makes some different trade offs but also gives a tight enough semantics that memory safety exploits aren’t going to work. Both CHERI and Fil-C are more comprehensive than Rust, since they attack the problem at the ABI level (and so you don’t get the problem that the protection only applies to the parts that were rewritten in the safe subset of a new language). Rust could be claimed to be better in that its compile time, but that doesn’t make a significant difference if you’re worried about the definedness of semantics or exploitability.

    • afdbcreid 7 hours ago ago

      Both only attach provenance to allocations. The common example is:

          struct User {
              char name[100];
              bool is_admin;
          };
      
      Where a buffer overflow can still overwrite `is_admin`.

      Both also require recompilation of everything, which might be possible for CHERI but not for Fil-C - which is why, for example, there cannot be Fil-C support for Windows or macOS.

      • aw1621107 5 hours ago ago

        IIRC Filip said that Fil-C can be modified fairly easily to catch those overflows, but that breaks a fair bit of C code.

      • cperciva 6 hours ago ago

        No, CHERI supports sub-object capabilities.

        • afdbcreid 6 hours ago ago

          Nice, didn't know that!

          • dwattttt 5 hours ago ago

            It's a safe/stable operation to derive a more limited capability (pointer) from a more capable one. An example would be an arena allocator having a pointer valid for the whole arena, but then only handing out a derived pointer that's valid for the size requested.

            You can also limit access permissions IIUC; I'm only going by old memory here, but you'd be able to hand out a read-only pointer derived from an RWX pointer.

  • vbezhenar 6 hours ago ago

    My main issue with UB in C is that it's silent.

    I'm OK with compiler doing wild thing, OK, whatever. Well, I'm not OK but I can accept it in this crazy world.

    But I want to have loud warnings! Like WARNING: this conditional operator has been collapsed to one branch because earlier division by zero is UB. And now I can notice it and rewrite it or just remove that condition.

    I understand that this code can be result of macro expansion. That's OK. Macros should either include some pragmas to temporary disable specific diagnostics or user should surround macro usage with these pragmas, if they can't edit the macro. It's already happening with other warnings.

    Or maybe compiler could be smart enough to distinguish macro expansion from honest user mistake, I don't know.

    I remember when C++ compiler just removed function epilogue where I wrote simple infinite loop. That was so crazy. So instead of entering the infinite loop, my program just continued to execute the function that happened to be linked below. Imagine debugging that. Zero diagnostics.

    • jakobnissen 6 hours ago ago

      This is completely infeasible. A C compiler makes so many assumptions that UB does not occur that every program, even well functioning ones, would emit walls of warnings. Also, these warnings can't be silenced. Consider the integer division case: Should the warning be emitted every time division by zero is encountered at runtime? In that case, C would be much, much slower. Or should the compiler emit the warning whenever it can't prove at compile time that the divisor is not zero? In that case, you would get tonnes of situations where you can't fix the issue because you can't prove to the compiler it's not zero. Division by zero is a easy case. It gets much harder to emit warnings for aliasing assumptions, for example.

      • bigstrat2003 6 hours ago ago

        > A C compiler makes so many assumptions that UB does not occur that every program, even well functioning ones, would emit walls of warnings.

        The problem here is that the C compiler ever assumes that UB doesn't happen. That is empirically very much not the case, therefore the compiler should never be allowed to assume a lack of UB unless it can somehow prove that to be true. I honestly don't really care how many optimizations that would break; correctness is king. Software that goes fast is only worthwhile if it works correctly.

        • dgrunwald 5 hours ago ago

          > I honestly don't really care how many optimizations that would break; correctness is king.

          It's approximately all optimizations. Good news: there's already a compiler option that does exactly what you want: -O0. It's even enabled by default (unless overridden by another -O switch)!

          • ahartmetz 4 hours ago ago

            -O0 code also omits a lot of "Don't be very stupid about it" optimizations, it's not a realistic option for much production code.

            • dgrunwald 4 hours ago ago

              A surprising number of those optimizations are only valid by "reasoning from undefined behavior".

              Under the as-is rule, optimizations must not change program behavior. A C program could theoretically use out-of-bounds pointers to scan its own stack, observing whether a value occurs on the stack. Thus, the as-is rule prohibits storing local variables in registers!

              But because out-of-bounds pointers are undefined behavior, the compiler can ignore the programs doing stack scanning, and so register allocation becomes possible under the as-is.

              So boring old register allocation is one of those "assume that UB doesn't happen" optimizations! Of course, the compiler never actually reasons "this pointer arithmetic is out of bounds therefore I can put that variable over there into a register" -- the reasoning from undefined behavior doesn't happen at compile-time, it already happened when the "register allocation" optimization was designed. But that's the case for most "assume that UB doesn't happen" optimizations! (this is also why it's so difficult for compilers to warn about undefined behavior -- they optimize based on it without ever detecting it!)

              If you want to eliminate "reasoning from undefined behavior", you'd also need to replace the "as-is" rule with something else -- an explicit list of allowed optimizations in the language standard?

              • afdbcreid 3 hours ago ago

                Indeed. No UB is something no mainstream C/C++ compiler does (and maybe ever did). Benchmarking the performance impact will require basically rewriting GCC or LLVM, an insurmountable task, because the assumptions are so ingrained.

                Though it might be possible to have more "sensible" UB. For example, assuming daemonic non-determinism for allocation address it is possible to turn stack storage into registers, without the need for an aliasing model. Benchmarking it has the same difficulty.

        • UncleMeat an hour ago ago

          > The problem here is that the C compiler ever assumes that UB doesn't happen.

          Imagine you have a tail-recursive function and you'd like to convert it to iteration. This is only semantic preserving if there is no UB!

          Or worse, if you can never assume that UB doesn't occur then you need all code to be compiled in a way that defends against data races. You can't change the allocation order on the stack, since there might be some write past the end of a buffer and now suddenly stack allocation order changes the semantics of the program. Nightmare.

          Imagine this program:

              int f() {
                int x = 1;
                a(x);
                b(x);
                return x;
              }
          
          Can we use constant propagation to replace "x" with 1 in the instructions? No. After all, a could smash the stack and write over x. You need a new read of that memory location every time you plan on using x.
      • vbezhenar 5 hours ago ago

        The compiler should emit a warning at compilation time, when it decided to replace the conditional with single branch, throwing away the condition itself and the second branch.

        If I wrote some code, I expect it to be present in the binary. I don't just write code to be removed by the compiler. If that expectation was wrong, compiler should inform me about that.

        • flohofwoe 5 hours ago ago

          > If I wrote some code, I expect it to be present in the binary.

          That's oversimplified. If after inlining and constant folding an if-condition turns out to be always true or false I would definitely expect that the compiler removes the dead branch.

          This type of optimization is the base for the fabled "zero-cost-abstraction" (which isn't only a C++ thing, C code depends on it just as much), and removing those optimization would seriously tank peformance in any non-trivial codebase.

        • cesarb 5 hours ago ago

          > If I wrote some code, I expect it to be present in the binary. I don't just write code to be removed by the compiler.

          It's very common to write code in templates or inline functions expecting the compiler to remove it if it's not relevant on the calling site. That's part of what makes "zero cost abstractions" have zero cost at runtime.

          For instance, I have a SIMD routine with extra code to process the tail (leftover elements smaller than the native vector size). When the compiler can prove that the size of the input will always be a multiple of the vector size (which is very common for my use cases), it will completely remove that tail handling code.

          • vbezhenar 5 hours ago ago

            They you'll see that warning and mark your extra code to remove that warning, because you're aware that it's subject to potential elimination.

            • microtonal 3 hours ago ago

              I think the point is that your code will light up as a Christmas tree, since basically anything could be eliminated in specific contexts. For example, the loop that handles vector-width parts of the array can also be eliminated (as in the condition/branching). If the compiler can prove that there will only be a small, fixed number of iterations, it will just unroll the loop.

        • quietbritishjim 5 hours ago ago

          How about if the compiler translates a pointer deference into a read of that memory address? That is potential undefined behaviour that has been "optimised" into something simpler than a safe operation (tracking all memory allocations and checking if the pointer is correctly pointing into one of them). So every pointer deference would also generate a warning (except perhaps where the compiler can prove from local information that it's safe).

          It is definitely a hopeless path.

        • UncleMeat 43 minutes ago ago

          Turning off all arithmetic simplification is a wild choice.

          Suppose you have the following program:

          unsigned f(unsigned x) { return 2*x; }

          Can this be optimized to instead perform a bit shift? That's faster than multiplying. But you wrote a multiplication operator. Should the compiler be responsible for keeping it?

    • flohofwoe 5 hours ago ago

      First we need to reduce the UB zoo because the types of UB in the C standard go all over the place (from "no newline at end of file" to "oops, this specific UB combined with this specific code compiled on this specific version of this specific compiler with these specific compile options leaks execution into the next function").

      Also AFAIK the point where UB causes 'runtime disruption' is way after the C frontend in the optimizer passes, e.g. much too late for issuing compilation warnings even if the UB situation could be detected (because as far as I understand the problem, the breakage happens mainly because of unexpected 'spooky actions at a distance' between different optimizer passes, e.g. a specific optimizer pass doesn't even notice that it broke the code).

      You can get runtime errors for a lot of serious UB problems via UBSAN though of course (at the cost of some performance).

    • masklinn 6 hours ago ago

      Most UBs are runtime conditions, and inserting runtime checks would defeat the point of optimizing for them not possibly happening.

      For the static ones there are often warnings you can set, but you’ll have to go through the list. Or possibly external checkers (e.g. clang-tidy has one for infinite loops but not sure it’s 1:1 with the optimiser on complex cases)

      • vbezhenar 6 hours ago ago

        It's not about inserting checks. It's about removing code that I wrote from the binary. That should not happen silently.

        • flohofwoe 5 hours ago ago

          You can get that behaviour (mostly) already today by simply not enabling optimizations.

          The resulting performance difference is basically the price to pay for such a 'strict' compiler which translates the input source code straight into machine code instructions without attempting to simplify the output code via inlining, constant folding and dead code removal.

          • vbezhenar 4 hours ago ago

            -O0 program is useless and should be written with any other language. The whole point of C is to use compiler with optimizations. I'm not against optimizations. I'm against optimizations that perform things unexpected by the programmer. Dead code removal is unexpected by the programmer, because programmer does not write dead code.

            • flohofwoe 4 hours ago ago

              Well you can't have your cake and eat it too ;)

              > Dead code removal is unexpected by the programmer, because programmer does not write dead code.

              Of course it is expected, because the optimizations leading to dead code are fundamental for the performance you get out of a release-mode binary (and that's also true for languages with less UB like Rust btw).

              When you have a function:

                  static int add(int a, int b) {
                      return a + b;
                  }
              
              And you call that function with const parameters:

                  int c = add(2, 3);
              
              Then you want that entire function call to be removed and "folded" into its result 5, and when this was the only place the function was called, you'd also want the actual function to be removed from the binary (because what's the point of lugging code around in the binary that's guaranteed to never be executed).
              • vbezhenar 3 hours ago ago

                The "code" is `a + b`. The rest is just decorations.

                Here's example from article:

                    static int f(int a, int b)
                    {
                     x = b ? 42 : 43;
                 return a/b;
                    }
                
                Compiler can inline `f` but if it removes conditional branch, that's where issue is. I wrote this branch because I expected `b == 0` to be a valid value. Compiler deduced that it can't be a valid value and decided to remove `43`. That's understandable. But let me know, so I'd rewrite my code myself instead. Because I have assumptions and compiler have assumptions and these assumptions do not match. It means that there's a bug.
                • masklinn 3 hours ago ago

                  So you expect the compiler to tell you every time it inlines a function with a null check in a function which already had a null check on the same value?

                  Because as far as it’s concerned that is the same thing.

                  Both clang-tidy and sonar will flag this since it’s broken on its own, but when you’ve fixed it you want the compiler to warn you every time this is inlined in a function where b is known non-zero and the check is redundant?

                  As flo told you, if you want a non-optimizing compiler use a non-optimizing compiler.

            • account42 3 hours ago ago

              I suggest you write such a compiler then or at least curate your own set of approved optimizations (compilers allow much more fine grained control than -O levels) if you refuse to take anyone's word that undefined behavior is required for good optimization.

              • masklinn 2 hours ago ago

                Strictly speaking it’s not, a better type system can provide intent as basis for optimisation e.g. rust’s unique references being definitionally non-aliasing does translate to optimisation opportunities (so much so that it found a ton of bugs in llvm’s noalias handling since it hit those in so many more ways than C or C++ codebases would).

        • amiga386 2 hours ago ago

          I agree but also disagree. You would be upset if it removed the condition from something like this:

              void clear_buf(char *buf, size_t len) {
                  char *end = &buf[len];
                  if (end < buf) return; /* in case len is so long the pointer wraps the memory space */
                  while (buf < end) *buf++ = 0;
              }
          
          (which it can do, by virtue of saying that constructing a pointer beyond end of buf is itself undefined behaviour, therefore it can expect that end is never less than buf. the more appropriate way to write this is "while (--len > 0) *buf++ = 0" and let the compiler invisibly change that to end=buf+len if it thinks that'll work in all situations and be more efficient than counting down len while moving along buf)

          ...but you wouldn't be upset if it changed this:

              for (int i = 0; i < 1000000; i++) {
                  if (i < 5) foo(i);
                  /*else something_else(i);*/
              }
          
          to effectively this:

              for (int i = 0; i < 5; i++) {
                  foo(i);
              }
          
          or even:

             foo(0); foo(1); foo(2); foo(3); foo(4);
        • UncleMeat 38 minutes ago ago

          Imagine this program:

            int main() {
              int x = 0;
              return g(x);
            }
          
            int g(int x) {
              if (x == 0) return 1;
              return bar();
            }
          
            int bar() {
               ...
            }
          
          Can this program be simplified to just

            int main() {
              return 1;
            }
          
          Simple inlining, constant propagation, and dead code elimination gets you there.
    • Gibbon1 3 hours ago ago

      ""code can be result of macro expansion"

      Might be wrong but the big issue with warning about UB probably isn't C it's C++ and it's templates. Modern C is mostly what you see is all there is. Where with C++ lord knows really.

      C is often a bit slower than C++. Which also means trading a little speed to eliminate UB isn't a big hit. Where doing that inside a template expansion can slow things down a lot.

      My preferred solution is people stop using C++.

      • gpderetta 28 minutes ago ago

        you are indeed wrong.

    • rramadass 6 hours ago ago

      Relevant:

      Memory error checking in C and C++: Comparing Sanitizers and Valgrind (quite comprehensive) - https://developers.redhat.com/blog/2021/05/05/memory-error-c...

    • UncleMeat an hour ago ago

      If your compiler concludes "you will encounter UB" then it will stop and yell at you. What you are encountering is something different.

      Suppose you have the following program.

          int f(int x, int y) {
            if (x < 0 || y < 0) return -1;
            if (x > 100 || y > 100) return -1;
            return x + y > 200;
          }
      
      A compiler class computes the possible ranges that x and y can take at the return statement. Another pass does arithmetic simplification and simplifies the final instruction to "return 0." Should the compiler warn you here? No, obviously.

      What you want is for this to only warning you if some conclusion about the program that led to a simplified conditional operation came from an assumption about UB. But compiler passes can't really track the "whys" of each conclusion it makes. It spreads so broadly so quickly.

  • hn_submit 8 hours ago ago

    I believe this is barking up the wrong tree since IMHO C is just "high level assembly" for systems programming. As soon as you add runtime behavior to combat Undefined Behavior (UB) you're blowing up execution times. And static analysis can only go so far without blowing up compile times.

    C is "the right tool for the right job" which is operating systems and its code which is called thousands of times per second. You cannot afford even one iota of runtime checks in that code. The developer must know what he's doing or he should get out of the kitchen.

    We should discourage the usage of C in application programming and prod application developers towards memory safe languages like Rust or Go.

    And I'm not even sure if Rust solves this case as far as UB is concerned.

    • SkiFire13 7 hours ago ago

      > C is just "high level assembly" for systems programming

      It is not, and that's the issue. If it was just "high level assembly" there would be no UB, and no need to have UB. Instead you have UB (although arguably not all of it is really needed) because you need optimizations, which in turn you need because otherwise C would be too slow for that "operating system" job.

      > code which is called thousands of times per second

      Scripting languages can easily have loops running thousand of times per second and even more. You're off by some order of magnitudes here if you want to describe operations that happen in operating systems.

    • gizmo686 8 hours ago ago

      A high level assembly would not have a UB problem, because the generated machine code would closely map to your source code, so even technically undefined behaviour would end up doing the expected thing for the given hardware.

      The problem with C is that modern compilers do a lot of transformations between your source code and the final machine code, so the actual behavior could be very far afield from what you would expect.

      > And I'm not even sure if Rust solves this case as far as UB is concerned.

      If your entire program is inside unsafe, then Rust is actually worse than C as far as UB is concerned. On the other hand, no one writes Rust like that, and Rust restricts all UB to unsafe blocks.

      • cesarb 5 hours ago ago

        > If your entire program is inside unsafe, then Rust is actually worse than C as far as UB is concerned. On the other hand, no one writes Rust like that, and Rust restricts all UB to unsafe blocks.

        Nitpicking: it's not the unsafe blocks themselves, but a few language features which are only allowed inside unsafe blocks. Language features which are allowed outside unsafe blocks work identically within unsafe blocks, and gain no extra UB just by being inside an unsafe block.

      • imtringued 4 hours ago ago

        Uh, no? If your entire program is inside unsafe then Rust is equal to C.

        Rust gets nasty if you want to build safe interfaces around unsafe code because you have to consider every single potential safe interaction (Rust aggressively reorders everything because the aliasing model allows it, in that respect it is "worse" than C but not less safe), whereas with an unsafe caller you're allowed to shrug with your shoulders and tell the higher up caller that it is his problem to figure out, but if you keep handing off to ever higher level unsafe callers, then you are literally back to C levels of unsafety.

        If it is not clear, if you use unsafe {} everywhere, you can use *const T, *mut T everywhere. Hence the no aliasing rules don't matter to you at all, which brings us back to the &mut T and &T unsafe-safe interactions between *mut T and *const T. Turning *mut T to &mut T or turning *const T to &T is hard.

        This is what makes it so hard to write unsafe code, it's supposed to be usable from a safe interface!

        It also explains why static mut was a huge mistake. You cannot take a mutable reference to them. Mutable reference semantics make no sense with static mut.

    • weinzierl 6 hours ago ago

      "And static analysis can only go so far without blowing up compile times."

      Compile time is not the main issue with static analysis. It's that C doesn't provide enough information and not the right information to make it efficient and effective.

      If you add this info you unavoidably will end up with something looking like Rust.

      (Whose long compile times are not caused by its static analysis parts BTW)

    • chasil 8 hours ago ago

      C is also famously bad at floating point optimization.

      Fortran has historcally led this realm (see the Numerical Recipes book).

      Julia is a newer option, and I understand that both are commonly used in Python objects.

      https://numerical.recipes/

      "Read the older 2nd ed. book in Fortran online for free."

      • jabl 7 hours ago ago

        There are several reasons why Fortran historically has been faster. Many of these are because Fortran is less strict about what the compiler can do so allows optimizations that wouldn't be allowed in C. Such as:

        - Procedure arguments are not allowed to alias. Similar to restrict pointers in C99+. This is often critical to allow loops to be vectorized, but the onus is on the programmer to ensure no aliasing or else you get UB.

        - Unspecified evaluation order for expressions. E.g. C requires that "a+b+c+d" be evaluated as "((a+b)+c)+d)" and with floating point it can't do it another way due to rounding. Fortran can do e.g. "(a+b) + (c+d)" where each subterm can be computed in parallel, but again at the cost of slightly different results due to rounding behavior for floats.

        - Old school Fortran lacked pointers which led programmers to program algorithms using arrays rather than fancier data structures, which cpu's love.

        In principle there's nothing preventing a competent C or C++ programmer can reach Fortran level performance. In practice, might be difficult.

        Of course, nowadays performance is much about designing for cache hierarchies (see e.g. "Data Oriented Design") where Fortran doesn't have a built-in advantage.

        • amiga386 2 hours ago ago

          I thought the main advantage had always been the row/column-major order of arrays:

          https://en.wikipedia.org/wiki/Row-_and_column-major_order

          In short, people solving numerical problems want to write like M[i,j] where i is iterated through on the innermost loop and j is iterated through on the outermost loop... and if the language works in sympathy with them and gets good cache-coherence, that's great, but if the language works against them and would make them write M[j,i] to get the same performance, they don't like it, so they keep writing M[i,j] and let the slowness be the other language's problem.

          • jabl an hour ago ago

            Nah, that is mostly about keeping the ordering in mind when iterating (e.g. do you iterate first over innermost or outermost index?). And there are algorithms such as (naive) matrix multiplication where the correct order for one matrix is the wrong for the other one, so plus ca change.

            But speaking of arrays, one definitive advantage of Fortran is that it has built-in multidimensional arrays in the language. Which is very nice when writing code that needs them. In C many people resort to the Fischer-Price My First Multidimensional Array implementation where they allocate each row separately, which is of course horrible. (In C++ you can use libraries like Eigen or Armadillo which use expression templates and operator overloading to have sensible multidimensional array support even if it isn't built-in.)

            • amiga386 37 minutes ago ago

              Uh... C does have multidimensional arrays. Writing "int a[5][10]" allocates the same contiguous storage as "int b[50]", and accessing a[3][7] will access the integer at ((char *)a) + 3*sizeof(int[10]) + 7*sizeof(int), which would be the same offset as b[37].

              It has to be row-oriented because the language allows you to take the address of each intermediate dimension, e.g. int *c = &a[3] gets you a pointer to the equivalent offset of b[30], and legally you can access c[0] through to c[9]. Accessing c[10] would be equivalent to a[4][0] and compilers will let you do it, but I think it's officially undefined behaviour?

      • mianos 7 hours ago ago

        There is not a lot of floating point in an operating system.

        The issue with floating point and aliasing preventing vectorisation was from the late 1980s when early C compilers lacked sophisticated alias analysis and standards were loose. is not a really a thing anymore. C can go as fast, specially compiled with strict aliasing. Maybe more work in compilation.

    • mdspan 7 hours ago ago

      Operating systems code can definitely afford runtime checking if it's not too expensive in terms of performance. The Linux kernel has a WARN_ON macro that does exactly this. The tradeoff is worth it in a lot of cases if you're exchanging a small amount of performance for greater debug-ability or security.

    • flohofwoe 4 hours ago ago

      > "high level assembly"

      That's already a fundamentally wrong assumption :) With today's compilers, C is a high level language like all the others. The optimizations happening to C code are not fundamentally different than for any other compiled high level language.

      > We should discourage the usage of C in application programming and prod application developers towards memory safe languages like Rust or Go.

      No that's rubbish, just as it would be rubbish trying to 'discourage' people from writing programs in assembly code or any other programming language.

      Ultimately, guaranteeing "safety" is the job of the sandbox your untrusted code is running in (e.g. the browser, operating system or VM).

      The only difference between Rust and any "unsafe" language should be that one fails already at compilation time and the other at runtime (because the sandbox killed your rogue process).

      E.g. if an operating system is exploitable because it allows untrusted programs to leak out of the sandbox, then that problem must be fixed in the operating system.

    • jandrewrogers 7 hours ago ago

      C is not “high level assembly”. You still need a model of what the code will create. This does not naturally follow from C code. Modern C++ is the most powerful systems language if you care about performance.

      While I don’t condone writing new code in C, there are major conceptions about its relationship to performance and C++.

    • imtringued 4 hours ago ago

      Hot take: The vast majority of UB optimizations in C are just hacks to get around mutable aliasing being the default.

      Edit: To be fair for the current topic it's more of an "Inverted Design-by-Contract" issue. The compiler infers the contract via UB instead of exposing it to the developer.

      when you divide a/b the contract b != 0 is formed on the function signature, but never made visible. Then someone inevitably violates the invisible contract and all hell breaks loose.

          // What the developer thinks the signature is
          int f(int a, int b);
      
          // What the compiler secretly changed the signature to
          int f(int a, int b) 
              __attribute__((requires(b != 0))); // The invisible contract!
  • bugake 41 minutes ago ago

    The fastest algorithms I can make for many problems use bit hacking. It works in x86, but it's technically undefined behavior.

    • davemp 19 minutes ago ago

      Often that’s just implementation defined behavior which is perfectly safe. You should just put some compilation guards to error out if someone tries to port your code.

  • 1vuio0pswjnm7 9 hours ago ago

    1790620504 | Reducing undefined behavior in the C language | https://lwn.net/SubscriberLink/1095811/b9325731ea9b61e0/ | https://news.ycombinator.com/item?id=49882419 | 15 comments

    1790673092 | Reducing undefined behavior in the C language | https://lwn.net/SubscriberLink/1095811/efcdbcf080cfa4c6/ | https://news.ycombinator.com/item?id=49890290 | 0 comments

  • rwmj 3 hours ago ago

    I think https://en.wikipedia.org/wiki/CompCert should at least be mentioned, even though it's not free software ("source available" license). It's the actual way that companies that have large C code bases in safety-critical industries verify their code.

    • afdbcreid 3 hours ago ago

      It not that related. It prevents miscompilations (and does reduce the amount of exploited UB) but does not make programs UB-free.

  • vrighter 3 hours ago ago

    i was actually thinking of starting a joke project which would make UB actually UB. not the sort of "in practice signed overflow is all done the same by most modern cpus". i mean an RNG and a compiler plugin system for introducing new Bs to go with the Us

  • elais-dev 10 hours ago ago

    i've seen static analyzers catch many ub patterns, but guaranteeing zero ub needs whole‑program analysis that blows up compile time and still produces false positives that drown developers

    • eru 9 hours ago ago

      Yes, it's pretty much impossible to prove anything about arbitrary programs.

      However if you are willing to restrict what programs you allow, you can make guarantees possible.

      Silly example: if you compile valid (safe) Rust programs to C, you know that the resulting code will not invalidate Rust's borrowing rules by construction; and in principle you could try to establish this guarantee just from the C code alone, never having seen the Rust original.

      However, you still wouldn't be able to have an algorithm that tells you for any arbitrary C code whether it has these problems or not.

      • robotresearcher 8 hours ago ago

        > if you compile valid (safe) Rust programs to C, you know that the resulting code will not invalidate Rust's borrowing rules

        If you know the compiler is correct, which you don't.

        • eru 3 hours ago ago

          For the illustrative Rust to C example, yes. But we have source languages for which we know the compiler is correct.

  • p1necone 10 hours ago ago

    The concept of undefined behaviour specific to C/C++ has always seemed batshit insane to me, and I'm yet to read anything about it that has made it seem any less so.

    • nananana9 9 hours ago ago

      It doesn't make sense to define what happens when you e.g. read from NULL because it's hardware specific - if you have virtual memory of some sort, you'd probably get a page mapping error. If you don't (embedded, WASM), you'd read back whatever value is at that address.

      Does Rust define what I get when I dereference NULL in unsafe code? I doubt it, since it would require a NULL check before every pointer dereference.

      The only insane thing about UB is that compiler writers took what everyone understand meant "the compiler emits what it emits and you get what you get" and turned it into "since it's undefined it means it can never happen so we can delete your null check".

      • tczMUFlmoNk 9 hours ago ago

        I don't disagree with your general thesis, but I don't think it's right to say that defining behaviors like dereferencing a null would require a null check before every dereference. For example, Java defines the behavior of dereferencing null—it throws a `NullPointerException`. My understanding is that JVMs implement this by (a) representing Java `null` references as the zero pointer, (b) mapping the zero page with write permission disabled, so that accesses are guaranteed to segfault, and (c) trapping `SIGSEGV` to translate that back into a Java exception.

        (Here's a link with a bit more detail: https://courses.cs.vt.edu/cs3214/spring2026/questions/catchi...)

        So, while it's accurate to say that inserting null checks before every dereference is one way that you could implement this to make it well-defined, that is not the only way. We have lots of clever tricks to solve problems more efficiently than may seem possible at first glance—Fil-C is a bit of a modern marvel in that regard!

        • nananana9 8 hours ago ago

          > My understanding is that JVMs implement this by (a) representing Java `null` references as the zero pointer, (b) mapping the zero page with write permission disabled

          If you want to run everywhere where C does, you can't rely on that. I gave WASM as an example - that's a widely used target that just exposes a flat memory model where 0 is literally just a normal address and there's no way to trap it (unless they've released extensions I'm unaware of). Same deal with most microcontrollers as far as I'm aware of, although I don't do embedded.

          I can't think of a way you'd implement null trapping efficiently on those platforms.

        • charleslmunger 9 hours ago ago

          Sure. But consider what would happen if you had an array, and you looked up the nth element. If that base address is a null pointer and the index is greater than the size of your zero page reservation, it'll get some other address which is holding stuff. There's ways to deal with that too, of course, but not for free and it carries implications for other things. In Java this is avoided because arrays carry their length at the beginning, and you check that first for bounds, so if the array pointer was null you'd fault a small number of bytes past 0 and it still works.

          Fil-C is amazing and a prime example that undefined behavior means implementor freedom, and the implementor can choose to always trap on null pointer use. Sometimes the implementor freedom doesn't buy you much; for example why should it be UB to do

              (const char*)NULL + 1
          
          Dereferencing null is and should be UB but why is just calculating a pointer problematic? I just did some research and some old architectures would actually trap on creating an invalid address. So if we want C to support those machines, the standard can't define the behavior to do something other than what the hardware does.
          • eru 8 hours ago ago

            That's another argument in favour of implementation defined behaviour, not undefined behaviour.

            Btw, a pointer in C doesn't necessarily need to mean an address (invalid or not) in your underlying machine. C is a formally defined abstract language, not portable assembly.

      • eru 9 hours ago ago

        > It doesn't make sense to define what happens when you e.g. read from NULL because it's hardware specific - if you have virtual memory of some sort, you'd probably get a page mapping error. If you don't (embedded, WASM), you'd read back whatever value is at that address.

        That's an argument in favour of 'implementation defined behaviour'. Not 'undefined behaviour'.

      • account42 2 hours ago ago

        > The only insane thing about UB is that compiler writers took what everyone understand meant "the compiler emits what it emits and you get what you get" and turned it into "since it's undefined it means it can never happen so we can delete your null check".

        That's just compilers optimizing your code under the assumption that it correctly follows the rules of the language, which is the basis for any optimizing compiler and completely sane. The compiler isn't optmizing any null checks that actually correctly prevent all UB only ones that are faulty and never fully worked to begin with and as a result are indistinguishable form any other pointless code paths.

      • afdbcreid 8 hours ago ago

        > The only insane thing about UB is that compiler writers took what everyone understand meant "the compiler emits what it emits and you get what you get" and turned it into "since it's undefined it means it can never happen so we can delete your null check".

        IMO UB as "undefined but don't be crazy please" was the original meaning of the standard but people argue on that. It is a fact that compilers didn't exploit UB as strongly back then. However there is a good reason for this change: if you want formal semantics (which you do want, at least possibly) it is pretty much impossible to distinguish the two. If "undefined behavior" is undefined in the math sense, or in formal semantics of languages - the operation can reach any Abstract Machine state, then the fact that you cannot reason about anything follows immediately. The only dubious thing is time-travel, and this was indeed removed in the last version of the standard (and also for Rust now).

      • p1necone 5 hours ago ago

        > The only insane thing about UB is that compiler writers took what everyone understand meant "the compiler emits what it emits and you get what you get" and turned it into "since it's undefined it means it can never happen so we can delete your null check".

        This is the bit that seems insane to me. Some of the behaviour of C/C++ compilers when they encounter undefined behaviour seems less like 'this is undefined, we'll just do something vaguely reasonable given the context/produce an error' and more like 'ahaha, the user has fallen into our trap, lets fuck them up'.

    • Veserv 9 hours ago ago

      Undefined behavior is just what happens when you violate a mandatory precondition.

      if (x > 0) {…}. But what if you entered the body when x <= 0?

      if (false) {…}. But what if you execute the body?

      These are “impossible”. What happens when the impossible occurs is “undefined”.

      When the older standards said signed integer overflow for addition is undefined what they are actually saying is that the real definition of + is:

          int +(int x, int y) { 
            assert(in_range(actual_math_add(x, y), signed_int_min, signed_int_max));
            return machine_add(x, y);
          }
      
      So of course what happens when you get signed overflow is undefined; you should hit that assert and your program should explode and die. You should “never” get to the next instruction.

      But, in the interest of performance, “release mode” (which in this case is just any compilation) elides asserts since as a programmer you should not write code that asserts in much the same way that you should not write assert(false) in a normal code path that is supposed to run. Assertions are intended for “impossible” code paths and usually get compiled out in “release mode” though maybe your code is buggy and can actually hit them and then your program goes off the rails because it had a bug.

      Put another way, if you did write assert(false) in a regular code path, would you find it unreasonable for the compiler to just delete the code after it? That is what undefined behavior is for.

      • drdexebtjl 8 hours ago ago

        This isn’t a good mental model. The compiler isn’t limited to just affecting the code after the undefined behavior. It is allowed to assume undefined behavior never happens for the entire program.

        For example, this code with an improper guard:

            if (!p) puts("error");
            printf("%d", *p);
        
        Since the program dereferences p in line 2, and dereferencing null is UB, the compiler is allowed to assume p is never null, so it’s allowed to delete line 1, even though it would have executed before the point where UB would happen.

        Even worse, the compiler isn't just allowed to not do things you told it to do, it's also allowed to do anything too.

        • dgrunwald 6 hours ago ago

          The compiler is only allowed to do that when it can prove `puts` will return (i.e. it must prove `puts` does not call `exit`). In a world with SIGPIPE that shouldn't happen. (unless maybe the compiler can prove that there's enough space available in the stdlib IO buffers so that puts won't actually output anything. but then there's no obversable difference in behavior)

        • account42 2 hours ago ago

          That's not worse but good. I want my compiler to not limit the performance of my code in service of impossible code paths and states.

        • Veserv 7 hours ago ago

          The question was: “The concept of undefined behaviour specific to C/C++ has always seemed batshit insane to me”.

          I was explaining why undefined behavior as a concept is a very sensible idea. Whether the expansive interpretation of the optimizations you are allowed to do when encountering the “impossible” are reasonable is a different question.

          • p1necone 5 hours ago ago

            Hence the 'specific to C/C++' part of that sentence. 'Undefined behaviour' is a probably unavoidable part of programming language design. The way that C/C++ interpret/handle it is batshit insane.

          • drdexebtjl 7 hours ago ago

            The part that is batshit insane is precisely that it makes the entire program impossible to reason about.

            If UB meant what you described, it would be sensible, but it doesn't, and it's not.

      • cesarb 5 hours ago ago

        > if (x > 0) {…}. But what if you entered the body when x <= 0?

        > if (false) {…}. But what if you execute the body?

        > These are “impossible”. What happens when the impossible occurs is “undefined”.

        By the way, that's exactly what happens with Spectre and its class of ghostly vulnerabilities! The processor speculatively executes these "impossible" paths, and discards the result once it detects they couldn't happen; but there are ways to "leak" information from that irreal world through side-effects like cache lines being discarded.

      • rramadass 7 hours ago ago

        Beautifully explained!

        UB is a formal tool for the optimizer.

        Most people who argue about UB on HN have no clue wth it means and how it differs from Unspecified and Implementation-defined behaviours. The standard already explains expressions/statements and how they relate to sequence-points/sequenced-before/sequenced-after code points which is what is needed to understand the anomalous behaviour above.

        Add in a introductory class in numerical analysis w.r.t. accuracy/precision/limits/rounding and the C/C++ programmer has enough knowledge to avoid problems in practice.

    • mpyne 9 hours ago ago

      It's not specific to C or C++, though it is more prominent there.

      Rust's unsafe mode, for instance, has undefined behavior. A whole list of them, in fact.

    • GrantMoyer 8 hours ago ago

      Let's take a simple example: writing past the end of an array. Allow me to argue with myself for a moment.

      > Surely the compiler can just check if each access is valid.

      Well, sometimes it can, but sometimes it doesn't know how long the array is. What if the array is passed as a pointer?

      > Maybe each array could be annotated with its size at runtime, and accesses could be checked at runtime too.

      That works, but it adds runtime cost that may legitimately be too much for some applications, for example, a Gameboy game (set aside that many Gameboy games were written in assembly).

      > Fine, so we'll make the programmer promise to ensure array accesses are always valid. Maybe they'll make a mistake sometimes, but what's the worst that could happen? Throwing your hands in the air and saying the compiler is allowed to do anything, that's just stupid.

      Well, maybe it's stupid, but this is one thing that could happen if you accidentally write past the end of an array: https://www.youtube.com/watch?v=Vjm8P8utT5g. I'm sure neither the programmers nor compiler writers intended that.

      Ultimately, the compiler can't guarantee any behavior if its assumptions are violated. The example may seem contrived, but it demonstrates that, given the right circumstances, the results of the logical contradiction are unbounded. This is a direct consequence of the "Principle of explosion": https://en.wikipedia.org/wiki/Principle_of_explosion. On second thought, maybe the runtime costs of array bounds checking are an acceptable trade-off after all.

    • chasil 9 hours ago ago

      It was written for a PDP-11 with 64k of RAM.

      There wasn't room for safe programming practices, and direct manipulation of the hardware was a design requirement.

      It assumes that you know what you are doing.

      There are also ports to the Zilog Z80, an architecture with similar limitations (UZI, FUZIX).

    • UncleMeat 20 minutes ago ago

      Imagine I have the following program:

          int main() {
            int x = 1;
            a(x);
            b(x);
            return x;
          }
      
      Can I use constant propagation to replace the reads of x with the constant 1? This relies on the assumption that there is no UB, since otherwise our function calls could smash the stack and overwrite where x is stored on the stack.
  • strenholme 8 hours ago ago

    I recently had a heated but productive discussion about undefined behavior here.

    As per the linked article:

    “There are currently about 100 instances of undefined behavior in the C standard, but the in-progress C2y draft has removed 45 of them.”

    I wonder how they handle the specific case of uninitialized but allocated memory.

    Let’s look at something which will result in undefined behavior in C99: [1]

      #include<stdio.h>
      #include<stdint.h>
      #include<stdlib.h>
      #define b(z) for(c=0;c<z;c++)
      uint32_t c,e[42],f[42],g=19,h
      =13,n[45],i,j,k;void m(){j=0;
      b(12)f[c+c%3*h]^=e[c+1];b(g){
      i=c*7%g;k=e[i++];k^=e[i%g]|~e
      [(i+1)%g];j=j+c;n[c]=n[c+g]=k
      >>j%32|k<<-j%32;}for(i=39;i--
      ;f[i+1]=f[i])e[i]=n[i]^n[i+1]
      ^n[i+4];b(3)e[c+h]^=f[c*h]=f[
      c*h+h];*e^=1;}int main(int c,
      char**v){char*q=malloc(2);if(
      q==0)return 0;q[0]&=31;q[0]|=
      64;q[1]=0;for(;;m()){b(3){
      for(j=0;j<4;){f[c*h]^=k=(*q?
      255&*q:1)<<8*j++;e[c+16]^=k;
      if(!*q++){b(18)m();b(2){j=c;
      b(4)printf("%02x",(e[1+j%2]
      >>8*c)&255);c=j;if(c%2)m();}
      puts("");return 0;}}}}}
    
    The key part of the above brick of code is this:

      char *q=malloc(2);
      if(q==0)return 0;
      q[0]&=31;
      q[0]|=64;
      q[1]=0;
    
    Here, we see that q[0] is an allocated but undefined byte. As per C99, this results in undefined behavior, however 20 years ago this was a good trick to get kinda-randomish bytes to use as a possible entropy source.

    Someone claimed that the above brick of code will compile in newer versions of clang such that, since the complex cryptographic pseudo random number generator code depends on uninitialized but allocated memory, the entire cryptographic operation isn’t performed.

    So I tested it against multiple versions of GCC and clang; I also tested it against TCC for good measure.

    In all cases, with all levels of optimization, the cryptographic routine ran. I even ran it against clang 23. In cygwin, it was a randomish but consistent byte (except for clang at a higher level of optimization, at which point the uninitialized byte had a value of 0); in Ubuntu 26, the uninitialized memory consistently had a value of 0 (in tcc/gcc/clang).

    I am hoping the up and coming C2y spec has very clearly defined behavior when using unintialized memory (ideally where it will work but the bytes can have any values).

    Naturally, I have updated my code to no longer use uninitialized memory as a source of entropy. 20 years ago, MacOS didn’t support clock_gettime() with nanosecond resolution, so that wasn’t a portable way to get pseudo-random bits; these days clock_gettime() is universal across modern development environments, and it provides pretty good entropy (along with using /dev/urandom in *NIX, which isn’t in POSIX but is widely supported, as well as CryptGenRandom() in the legacy Win32 port). [2]

    [1] Said person said the appendices to C99 aren’t authoritative, but if something is in the spec, including in the appendices, it’s authoritative.

    [2] I don’t blindly trust /dev/urandom to always make really hard to guess pseudo-random bits, because my code is open source, and, as such, doesn’t just compile in Linux. It often times will be compiled in embedded systems, and even Linux has had at times issues with /dev/urandom on Raspberry Pis.

    [3] I would also like to see uint8_t, int8_t, uint16_t, int16_t, uint32_t, int32_t, uint64_t, and int64_t mandated. They exist in C99, but aren’t mandated, even though every real world compiler from this century supports all of the above types. Yes, I know about _BitInt(8/16/32/64/128/etc.) but a compiler from 2004—and yes I still use one to make win32 binaries—doesn’t support these new C23 datatypes.

    • jcranmer 7 hours ago ago

      > [1] Said person said the appendices to C99 aren’t authoritative, but if something is in the spec, including in the appendices, it’s authoritative.

      That's not true. There is a difference between normative text and informative text. Informative text is not authoritative, and you were citing an appendix that is labeled as informative.

      > I am hoping the up and coming C2y spec has very clearly defined behavior when using unintialized memory (ideally where it will work but the bytes can have any values).

      It won't. Uninitialized memory can't have "very clearly defined behavior" without breaking essentially every single implementation, and WG14 is very loth to break existing implementations.

      • afdbcreid 7 hours ago ago

        > It won't. Uninitialized memory can't have "very clearly defined behavior" without breaking essentially every single implementation, and WG14 is very loth to break existing implementations.

        You can initialize with some bit pattern. This is what was accepted for C++ (but only for certain types of memory).

        • vbezhenar 6 hours ago ago

          That's huge performance loss. Imagine doing useless initialization of multi-megabyte chunk of memory. Not acceptable for C programs.

          In fact I hate that C mandates static variables being initialized to zero. This is dumb. When I'm writing for MCU, that's useless cycles spent at the power on. Thankfully it's possible to fix with linker tricks, but it should not be an issue in the first place.

          • aw1621107 6 hours ago ago

            In the case of C++ initialization is (currently) only done for automatic variables, so you're not likely to be zero-initializing multiple megabytes of stack. There's also [[indeterminate]] for opting out of initialization, though that has to be done per variable.

          • afdbcreid 5 hours ago ago

            I agree it is, and like the other commenter said even C++ only did it for some things. I was only saying it is not breaking.

      • rallyrascal 5 hours ago ago

        How would it break existing implementations?

        Having accessing uninitialized memory be UB is just insane, there's only two possible sane implementations - potential trap for caps based systems (or for a static analyzer) or you get a pseudo random value i.e 'whatever happened to be there'. So just make it implementation defined.

    • afdbcreid 7 hours ago ago

      It is quite hard to come with a reliable example, but (https://c.godbolt.org/z/56bfYszEb):

          int* p = malloc(sizeof(int));
          int v = *p;
          if (!(v < 0 || v == 0 || v > 0)) {
              exit(1);
          }
      
      This compiles into `exit(1)` in Clang under -O3.
      • kstrauser 6 hours ago ago

        How? That doesn’t seem like it should be possible. Why is that?

        • robinsonb5 6 hours ago ago

          As I understand it, the compiler "knows" the range of valid values for a variable (for instance, if an int was promoted from an unsigned char, it knows it doesn't need to bother with testing for negative values). If the value came from uninitialised malloc'ed memory, the compiler knows that there are no valid values less than zero, no valid values directly equal to zero, and no valid values greater than zero - thus all three tests fail, and the overall negated test always succeeds!

          • afdbcreid 5 hours ago ago

            The reasoning is roughly correct but range checking isn't what const-folding the comparisons into `false` (rather it's a special rule for LLVM's `undef`).

        • afdbcreid 6 hours ago ago

          Because it invokes UB.

    • fragmede 7 hours ago ago

      Interesting. x86_64 has RDRAND for entropy (and RDSEED to seed). Oh I guess there's also ARM these days but I bet they got one too.

      • strenholme 7 hours ago ago

        I write code which runs in embedded spaces, so a lot of ARM and even RISC-V. So I have a simple, portable source of pseudo-random numbers which will give good random numbers on a potato, which means rolling my own.

  • fithisux 7 hours ago ago

    We need to learn to use assembly and call it from C. C was never meant to be the end of programming. Reducing undefined behavior is the least the standards body can do. Fil-C is also an extremely useful tool.

    The real problem I see is the inability of independent compiler writers upgrading to the latest standard. I think this is also an issue that the standards body should pay attention. Help implementers.

  • BeaverGoose 10 hours ago ago

    Make signed overflow defined please.

    • _kst_ 9 hours ago ago

      I disagree, though I wouldn't mind adding a mechanism to say that you want signed overflow to be well defined.

      C23 already requires 2's-complement representation for signed integer types, but signed overflow still has undefined behavior. I think that mandating 2's-complement wraparound would be a mistake.

      Some instances of undefined behavior can be detected at compile time. For example, if I write

          int too_big = INT_MAX + 1;
      
      a reasonably clever compiler can warn about it (and in fact both gcc and clang do so). If the result of INT_MAX + 1 were defined by the language to be INT_MIN, there would be no basis for such a warning.

      If you evaluate n + 1 and it's possible for n to be equal to INT_MAX before the addition what do you want the result to be? Would quietly yielding INT_MIN really be useful?

      Ideally, if I (accidentally) evaluate INT_MAX + 1, I'd like to be told that I've made a mistake. C doesn't have a good mechanism for doing so.

      gcc has a non-standard option "-fsanitize=signed-integer-overflow" that can be used to catch signed overflow at runtime. If signed overflow yielded a well defined result, that option would be non-conforming.

      • eru 9 hours ago ago

        > a reasonably clever compiler can warn about it (and in fact both gcc and clang do so). If the result of INT_MAX + 1 were defined by the language to be INT_MIN, there would be no basis for such a warning.

        Compilers warn about perfectly well defined behaviour all the time. That's why these are warnings, not errors.

      • afdbcreid 8 hours ago ago

        You can perfectly specify that integer overflow either traps or wraps-around (this is what Rust does). Then the flag would be conforming. You can also say it is Erroneous Behavior (a new term in C++, not yet adopted for C AFAIK) and specified to wrap-around, which will also allow the flag (and arguably models Rust behavior more closely).

    • account42 2 hours ago ago

      I would rather make unsigned overflow undefined by default and and specific defined overflow types or builtins.

      Almost always, overflow is a logic error no matter what behavior the compiler defines for it so you need to make sure it doesn't happen anyway.

    • creato 9 hours ago ago

      Why would this help? Unsigned integer overflow is defined behavior, that causes basically the same set of bugs in practice. In a way it is worse, because at least runtime UB checkers have a reason to complain about signed integer overflow, but they won't complain about unsigned integer overflow.

      • rallyrascal 4 hours ago ago

        If it's well defined you can actually write simple code to check for overflows (like check if the sum or product is less than the inputs) and it wont get deleted by the compiler.

        • saagarjha 2 hours ago ago

          Consider using the simpler code in your standard library to do this operation instead of the less scrutable version you've described here.

    • eru 9 hours ago ago

      It would be an improvement.

      However if you wan, you can already get that via a flag in pretty much any C compiler you care about.

    • homosapien97 8 hours ago ago

      Use -fwrapv

      • pajko 8 hours ago ago

        Or -ftrapv depending on usage, both converting UB to IB. Generally letting a number overflowing and wrapping around is opening the door to let in the bad things creeping in the shadow.

  • jdw64 10 hours ago ago

    If we follow this video, the real constraints essentially mean POSIX and ABI. This implies that the contract a programmer must understand is distributed across multiple layers—not just the language specification, but effectively the operating system and ABI as well. Ultimately, this leads to the conclusion that even if the language itself is fundamentally free, constraints are inherently necessary at its lower layers. If we were to bloat the compiler—that is, if we restricted freedom like Rust does with its borrow checker—then the freedom available to the programmer would vanish. If that happens, people might grow weary of a language that is supposed to be free. In the end, any single layer inherently restricts freedom. In other words, I feel there is a need to transfer the complexity that a programmer must manage over to a mechanical management system, but which layer would be best for that?

    Currently, based on experience, this level of complexity is categorized into the language layer, and that level of complexity into the operating system layer. But in the future, won't there be some sort of complexity theorem that determines which layer minimizes complexity the most, and won't systems be completely rewritten based on that?

  • nine_k 6 hours ago ago

    I frankly find the situation around UB puzzling. AFAICT the idea was born in the times of very anemic compilers which would translate things that don't make definite sense, or work unpredictably on certain architectures. But why is it still a thing now, 50 years later?

    If a compiler can detect UB, I think the only reasonable way for it is to have the program terminate immediately, not release "nasal demons". (And, of course, many bugs get smashed against the -Wall which should be the default.)

    Of course, not all UB can be detected at compile time, like reading the padding bytes in a struct mentioned in the article. I wish there was a way to opt out of all UB, in an arbitrary but predictable way (an immediate crash would be fine), something like -fsanitize but for any UB at all that might happen at runtime. In quite some important code, I'd agree to tolerate the performance hit, it may be cheaper than dealing with the aftermath or an RCE exploited. I see that it's not entirely realistic though.

    • masklinn 6 hours ago ago

      > If a compiler can detect UB

      In the general case it can not, that is why they were made UBs in the first place.

      And that is generally not how compilers see UB, they don’t look for UBs lying around, they assume UBs can’t happen and use that for constraints, which they then propagate.

      For instance

          *i = 1
          if (id) { … }
      
      The compiler will likely remove the check, because the dereference tags i as non-null, which makes the test redundant.

      And this occurs and is extremely desirable every time e.g. a function is inlined in an other one which already checked for null.

      • nine_k 6 hours ago ago

        But how is your example UB? It's straightforward data flow analysis.

        AFAICT, UBI is stuff like that:

          short int i;
          for (i = 0; i < 33000; i++) do_something();
        
        Whether this program continues as normal, crashes, or enters an infinite loop is platform-dependent, because short int can be as small as 16 bit, and an integer overflow past 32767 can be a crash or a wraparound.
        • masklinn 3 hours ago ago

          > But how is your example UB?

          The goal of my comment was to demonstrate how compilers use UBs to drive analysis.

          Dereferencing a null pointer is one of the most basic UBs you can find in C. Since a valid program can not contain UBs, the compiler concludes i must be non-null, propagates that to the test, which is constant and true, thus redundant and eliminated.

        • d-u-d-e 4 hours ago ago

          Dereferencing i is undefined behavior if it's null.

      • imtringued 3 hours ago ago

        "And this occurs and is extremely desirable every time e.g. a function is inlined in an other one which already checked for null."

        Again with the backwards word choices...

        C developer: "this is extremely desirable"

        Translation: Our system is unable to specify intent (nullability of the pointer) so we just guess.

        Further context: Other systems specify intent and have no need to guess and as a consequence have this "extremely desirable" feature. The "clever" C tricks are always told from a C centric perspective.

        • masklinn 3 hours ago ago

          There is nothing backwards, the exact same analysis can be and is done on other types. If I write

              if v.is_none() {
                  return 0;
              }
              v.unwrap()
          
          (or any other similar pattern) and enable optimisations, rustc will also elide the second check and the impossible panic call. Because any time you inline there’s good odds you’re creating redundant code (and other such).

          That is the main reason inlining is so fundamental to modern optimizing compilers, it’s nice that it removes funcall overhead, but it’s more relevant that it unlocks a whole slew of further optimization opportunities (including even more inlining).

    • peterfirefly 2 hours ago ago

      > in the times of very anemic compilers

      These days, it's mostly taken as an invitation by very powerful compilers to sometimes do heroic optimizations (that we like) -- and sometimes screw us over in really obscure ways (which we don't like). We would really like to keep their ability to perform heroic optimizations and that's why there's so much pushback against eliminating undefined behaviour.

    • pjmlp 2 hours ago ago

      There wasn't even the case, C compilers were anemic, not the ones being developed for PL/I, NEWP and all other derived systems languages.

      "Oh, it was quite a while ago. I kind of stopped when C came out. That was a big blow. We were making so much good progress on optimizations and transformations. We were getting rid of just one nice problem after another. When C came out, at one of the SIGPLAN compiler conferences, there was a debate between Steve Johnson from Bell Labs, who was supporting C, and one of our people, Bill Harrison, who was working on a project that I had at that time supporting automatic optimization...The nubbin of the debate was Steve's defense of not having to build optimizers anymore because the programmer would take care of it. That it was really a programmer's issue.... Seibel: Do you think C is a reasonable language if they had restricted its use to operating-system kernels? Allen: Oh, yeah. That would have been fine. And, in fact, you need to have something like that, something where experts can really fine-tune without big bottlenecks because those are key problems to solve. By 1960, we had a long list of amazing languages: Lisp, APL, Fortran, COBOL, Algol 60. These are higher-level than C. We have seriously regressed, since C developed. C has destroyed our ability to advance the state of the art in automatic optimization, automatic parallelization, automatic mapping of a high-level language to the machine. This is one of the reasons compilers are ... basically not taught much anymore in the colleges and universities."

      -- Fran Allen interview, Excerpted from: Peter Seibel. Coders at Work: Reflections on the Craft of Programming

    • vbezhenar 6 hours ago ago

      My understanding is that terminating program requires instructions and optimizing program involves removing "useless" instructions. If your code is littered with checks and jumps, it won't be faster than code without these checks and jumps.

      In other words: if you write `a[i]` in C, the compiler could either add index checks if its knows the size of `a` with `abort()` calls; or it can just compile it to single access instruction. The latter is faster. The former might be useful for debug build, I guess, but otherwise it's too slow to be acceptable for C usage.

      • nine_k 6 hours ago ago

        Yes. Carrying around stuff like ASan / TSan and a ton of other checks is just pretty expensive, and it does not even guarantee much. They only way is to use a better language %)

    • imtringued 4 hours ago ago

      As I already said C compilers use UB as an inverted design by contract system that is invisible to the programmer.

      You dereference a pointer? That changes the contract of the function so that you must never pass in a non-nullable pointer.

      Divide by b? You are not allowed to pass in a zero as b parameter to the function.

      But you don't get to see that. Nobody tells you.