Skip to content
View horecdev's full-sized avatar
  • Warszawa
  • 02:46 (UTC +02:00)

Block or report horecdev

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
horecdev/README.md

Michał Horecki

I like building AI type stuff from the ground up <3
Freshman CS student @ Warsaw University of Technology (EiTI).

Things I've built

  • GradCraft: a deep learning framework in C++/CUDA with an autograd engine, fused kernels and custom memory pools. Trained a 90M-parameter GPT on 1.5B tokens on a single RTX 3090, at 45-50% of PyTorch eager throughput.
  • Speech denoiser: an LSTM in plain NumPy, with hand-derived backprop through time (BPTT) and my own STFT/ISTFT. Audio samples are in the repo.
  • Lyme detection pipeline: a 3-stage TensorFlow vision pipeline, built as a paid project.

...and many more. This README is not a spreadsheet of everything I have ever coded.

Now: building an object tracker on a servo motor and esp32 cam.

🔗 📄 CV (1 page) · X (Twitter) · Email

Pinned Loading

  1. GradCraft GradCraft Public

    A C++ Autograd Engine written without ANY external frameworks. It can train a pretty big GPT locally. Built to maximize performance.

    C++ 1

  2. LSTM-denoiser-from-scratch LSTM-denoiser-from-scratch Public

    Long Short-Term Memory built without high-level ML libraries

    Python

  3. tiled-attn-fwd-3070ti tiled-attn-fwd-3070ti Public

    Tiled attention kernel vs PyTorch Naive vs PyTorch FA

    Cuda

  4. sgemm-3070ti sgemm-3070ti Public

    Naive vs Tiled vs cuBLAS GEMM benchmarked.

    Cuda