Parth Badgujar
  • About
  • Blog (current)
  • formatting
  • •

  • images
  • •

  • links
  • •

  • math
  • •

  • code
  • •

  • blockquotes
  • •

  • external-services
  • Beating torch.compile with Megakernels in CuTe DSL [Part 1]

    A deep dive into custom GPU kernels and how they can outperform torch.compile.

    43 min read   ·   July 05, 2026

    2026   ·   gpu   cuda   triton   machine-learning   pytorch   ·   tech

© Copyright 2026 Parth Badgujar. Powered by Jekyll with al-folio theme. Hosted by GitHub Pages. Photos from Unsplash.