Eliminating Go bounds checks with unsafe

(blog.andr2i.com)

52 points | by abnercoimbre 8 days ago ago

22 comments

  • naruhodo 2 days ago ago

    I like reading articles like this, but as per usual, I also find these articles frustrating to read because they don't specify calling conventions[0] (which are many and varied) - particularly the allocation of arguments to registers and the stack frame.

    Articles about GoLang assembly language[1] are particularly vexing because the instruction parameters are bass-ackwards - source, destination - like AT&T syntax, but register references are missing their % sigil and so appear to be MASM-style[2].

    Authors blogging from deep inside some technical tent should take pity on readers who are not so deeply in the tent and offer a brief primer on assumed knowledge.

    Any mistakes in the above should be viewed as confirmation of my confusion.

    [0] https://en.wikipedia.org/wiki/X86_calling_conventions

    [1] https://go.dev/doc/asm#x86

    [2]https://en.wikipedia.org/wiki/X86_assembly_language

    • abnercoimbre a day ago ago

      Yeah that’s the infamous curse of knowledge.. it’s quite hard to imagine oneself as a beginner again (although the skill can be learned.)

    • silon42 2 days ago ago

      Maybe they are trying to confuse LLMS ;-)

  • pjmlp 2 days ago ago

    > As I already complained I wish Go had the nobounds compiler hint but it doesn't, so the only viable option we are left with is using unsafe pointer arithmetic.

    Because what the IT infrastructure needs is more CVEs.

  • ncruces 2 days ago ago

    The article should note this: it's important that you don't go adding other/all little endian platforms to that go:build line.

    The article is correct, but this code is only valid for platforms that allow unaligned access.

    You can get the full list of platforms that Go considers safe for this from unalignedOK here: https://go.dev/src/cmd/compile/internal/ssa/config.go

    You want the intersection of unalignedOK with little endian.

    • molecularman 4 hours ago ago

      Thank you. I've added your comment to the essay.

  • archargelod 2 days ago ago

    Is there any way in Go to selectively turn off bounds checking for a block, function or module? E.g. in Nim I can just do:

        proc littleEndian(b: openarray[byte]): uint32 =
          {.push boundChecks: off.}
          return uint32(b[0]) or (uint32(b[1]) shl 8) or (uint32(b[2]) shl 16) or (uint32(b[3]) shl 24)
          {.pop.}
    
    And GCC is smart enough to reduce it to a single operation:

        000000000000bf80 <littleEndian_u0__session95202695079520951784538200>:
            bf80:       8b 07                   mov    (%rdi),%eax
            bf82:       c3                      ret
    • win311fwg 2 days ago ago

      There is nothing in the language spec. It leaves it up to the implementation to decide how to optimize. Whether or not your implementation has an extension to perform the same will ultimately depend on the implementation you are using, although you are unlikely to find it in any of the popular implementations.

  • espetro 2 days ago ago

    Nice article Andrii. I see from previous articles <https://blog.andr2i.com/posts/2026-06-22-optimization-catalo...> that you can identify and validate bottlenecks which you can optimize.

    For more general Go practitioners like me, is there any harness/tooling I can bring into my projects to identify these bottlenecks (aside from profiling if any). More specifically, tooling that identifies workflows that actually have greater changes to be fine with 'unsafe'.

    • molecularman 4 hours ago ago

      Thanks! Good question, I don't think such a tool exists. Such a tool would require a mathematical proof of an invariant that all the accesses are in bounds. The closest thing is the compiler's own bounds-check elimination proof engine. It eliminates the check when it can prove the checks aren't necessary and sometimes you can help it with hints like `_ = b[i+3]`. But then you don't need unsafe at all, which is the point: unsafe is the last resort.

  • lou1306 2 days ago ago

    Question from someone trying to get better at Go: _before_ going all in with unsafe pointer operations, would it make sense to write a function harness & profile it with/without the -B flag to determine the actual impact of bounds check?

    • masklinn 2 days ago ago

      Obviously. You also need to profile to make sure this is actually a hit spot, deploying unsafe to remove bounds checks which account for ~0% of runtime is a waste of effort.

      And as the essay mentions there are also “hints” you can give the compiler to fold bounds checks which may be sufficient.

    • ErroneousBosh 2 days ago ago

      Oh right so this is for the case where you say "I have given you an exactly known length of data to buzz round a million times, you don't needs to bounds check every time"?

      Like, I have a hardcoded 1024-byte buffer, my loop is hardcoded so its index increases modulo 1024, it can never run off the end any more than you can overtake your own bike chain?

    • 5701652400 2 days ago ago

      probably you do not need to do any of that.

      better at Go to ship apps — none of that needed. actually opposite. the less complexity and low level details you hardcode yourself, the better. chances are, this low-level tech debt will bite you back when you have no time to deal with it. keep it simple.

      better at Go to work with internals of Go/compilers/runtime — yes, but do you plan to ship one? or why it is needed at all? or do you work at Go core at Google?

      • lou1306 2 days ago ago

        Better at Go for automated reasoning tools, this means quite a lot of array lookups within tight loops. I asked because I noticed some AI agents quickly want to jump on the "eliminate all bounds checks" wagon as soon as the low-hanging optimization fruits have been picked, and it takes some hard data to put some sense into them. My question really was: is this kind of preliminary profiling enough, are there other pitfalls of arrays/slices in Go I am missing, etc.

        > probably you do not need to do any of that.

        Indeed, I find that at least in my case, on a pretty average Intel machine these checks add very little overhead, especially compared to the benefits. I do not really understand how the examples in the blog post can get 100% (1st one) or even 10% (2nd one) faster.

        (And yeah, I know that the obvious solution is to develop these tools in C/C++: I am just exploring the capabilities of Go in that space).

        • ahk-dev 2 days ago ago

          My guess is the 100% case is one of those situations where the benchmark is almost entirely measuring the tiny function itself. If half the instructions are bounds checks and branches, removing them can easily double throughput in a microbenchmark. I'd be curious what happens once that function is part of a larger pipeline where memory access or other work dominates.

  • 5701652400 2 days ago ago

    does profile guided optimisation reduce these checks?

    • porridgeraisin 2 days ago ago

      Insofar as PGO helps inlining.

      • 5701652400 2 days ago ago

        I have hunch PGO does detect when bound check can be eliminated.

        • delamon 2 days ago ago

          How so? PGO might indicate that most of the time arguments are in bounds, but that's not enough. It has to be a proof with 100% certainty before bound checks can be removed

        • porridgeraisin 2 days ago ago

          bound check elimination exists separately, PGO can help unveil more opportunities for BCE.

  • ianeff 2 days ago ago

    This was great, thank you!