Skip to content
View sleepyeldrazi's full-sized avatar
  • Hamburg, Germany
  • 01:29 (UTC +02:00)

Block or report sleepyeldrazi

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
sleepyeldrazi/README.md

Kaloyan Nikolov

Software engineer working on LLM inference and training. M.Sc. Computer Science @ RWTH Aachen.

Active work on git.kokoham.com.


Current focus

  • Multi-Token Prediction (MTP) and speculative decoding for local inference
  • KV cache quantization and Metal GPU kernel optimization
  • Diffusion-based training for hybrid attention+linear architectures research
  • Ternary weight quantization research

Active projects

ds4-nvfp4-spark — fork of ds4 adapted to run expert-prunned, mixed NVFP4 quantizations of DeepSeek V4 Flash

omlx — personal fork with MTP decoding and Q4 KV cache with Hadamard rotation

sleepy-llm — Zig-native LLM inference engine with hand-tuned Metal kernels

sleepy-agent — fully local Android AI assistant, on-device Gemma 4 inference


Background

  • Cross-compiled the core RASR ASR inference engine from x86 to ARM and Android.
  • Built a streaming on-device ASR demo on Pixel 6, benchmarking ONNX Runtime and TFLite as independent backends.
  • Trained streaming ASR models in PyTorch with causal topologies and hard latency constraints for edge deployment.

Stack

Python · Zig · C/C++ · TypeScript · Kotlin

Pinned Loading

  1. ds4-nvfp4-spark ds4-nvfp4-spark Public

    Mixed NVFP4 serving of DeepSeek V4 Flash on DGX Spark (GB10) - fork of antirez/ds4 with REAP expert pruning, NVFP4 quantization, FP8-packed KV cache, and managed-memory serving

    C 8

  2. sleepy-agent sleepy-agent Public

    Fully local AI assistant for Android with Gemma 4. Voice, image, text input. Web search capable.

    Kotlin

  3. little_helper_tui little_helper_tui Public

    little helper -- Spectre.Console TUI for LLM coding agents. Local-first, multi-provider (OpenAI/Anthropic), token budget, session history, diff viewer, model arena.

    C#

  4. aarch64-laptops/build aarch64-laptops/build Public

    Build an Linux OS based image

    ASL 268 68

  5. linux-surface/kernel linux-surface/kernel Public

    Linux kernel with modifications for Microsoft Surface devices.

    C 157 50