Skip to content
View sgoel2be24-cyber's full-sized avatar
💭
💭

Highlights

  • Pro

Block or report sgoel2be24-cyber

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
sgoel2be24-cyber/README.md

Shikhar Goel

Software Engineer · distributed systems, static analysis, native apps — and the AI layer on top

ranked #1 on the OSCI'26 leaderboard with 240+ merged PRs · won a special award at HackBlox 2026 · reached the DataForge 2026 finals at IIT Kharagpur · finished in the top 1% at three HackerRank Orchestrate editions · shipped a 48-operation Pinecone integration into Corsair · merged a file-descriptor leak fix into NumPy · landed two changes in Google's Agent Development Kit · landed three bug fixes in Apache Airflow · fixed the license-compliance audit in Apache Magpie · closed a CLI dry-run hole in Cognee · built a Go job queue that survives kill -9 with zero job loss · wrote a ReDoS analyzer that proves each finding with an attack string · built an incremental build engine with a content-addressed cache · built an offline screen assistant for blind users on the Snapdragon NPU · built an Android assistant you teach by doing a task once



Rankings and awards Selected builds Open source Systems and software AI engineering Experience Toolkit


whoami

Education B.E. Computer Engineering — Thapar Institute of Engineering & Technology, 2024–2028
Based in India
Works on Durable backend systems · static analysis · build and test tooling · native macOS and Android · on-device and agent AI
Currently Building Netra, an offline screen assistant for blind users on the Snapdragon NPU, and Teachable Voice, an Android assistant that learns a task from one demonstration
Reach me Portfolio · LinkedIn · shikhardeepgoel@gmail.com

I build durable backend systems and developer tooling — crash-safe storage, fault-tolerant dispatch, and analyzers that produce a working exploit rather than a warning. The agent and evaluation work runs on the same stack: deterministic checks around probabilistic models.


Rankings & awards

Result Event
🥇 #1 on the live leaderboard Open Source Connect India '26: 240+ merged pull requests across six projects. Leaderboard
🏆 Special Award Winner HackBlox 2026, open-source hackathon by Hackers Cult, judged on best project. Built Pact
🏅 Finalist DataForge 2026, IIT Kharagpur: Synapse Memory Lab
Top 1% in three editions HackerRank Orchestrate 2026, global 24-hour AI-agent hackathon

Selected builds

Each card replays the mechanism the project is built around. Click one to open the repo.

conveyor: jobs are appended to a write-ahead log, the broker is killed with kill -9, the log is replayed in 7.7 ms and no job is lost redoscope: a witness attack string grows while regex match time grows exponentially, and the blow-up is verified

forgegraph: a cold build runs all 11 tasks on three workers; after editing src/api.txt only the 4 affected tasks rerun and 7 come from cache teachable voice: a command is demonstrated once, then a new spoken value is replayed through the same taps, stopping at the payment screen


Open source

Open Source Connect India '26 — currently #1 on the live leaderboard, with 240+ merged pull requests across six projects.

What those pull requests fixed

Most landed in SecureFlow, a GitHub-App security scanner — broken Prisma queries, double-counted findings, mis-compiled ignore globs, rate-limited AI explanations, and worker DLQ routing coverage taken from 57% to 96.5% — and in Truxify, a logistics platform, re-landing reverted reliability fixes for retries, circuit breakers, upload scanning and sharding. The rest went to Air-Quality-Intelligence — timezone-correct UTC timestamps, shared in-memory SQLite connections, half-open detection windows and per-station outage handling — TCalc — git-status path matching, token-drift report totals and agent-rule generators — and adaptq, whose CMake build now works on macOS arm64.

Upstream merges, each linked to its merge or landing commit. Click a change to see what it fixed.

Project Status Change
apache/airflow #71691 · #73323 · #73324 ✅ merged
Three bug fixes, in Airflow core and the Amazon and Google providersThe deprecation shim in airflow.utils.helpers used bare __import__ on dotted paths and got the top-level module back instead of the submodule, so moved functions like render_template_as_native raised ImportError; switched to importlib.import_module, with regression tests (bb3e71d). With deferrable=True, the EMR Serverless delete operator's inherited stop step deferred before any delete ran, so the task resumed on the stop event and reported success with the application still alive; stop completion now goes to a handler that validates the event and then deletes, closing issue #72123 (c5d7f60). With psycopg 3, PostgresToGCSOperator's server-side cursor read rows through fetchone(), one FETCH FORWARD 1 round trip per row with cursor_itersize ignored, so a reported 2M-row export slowed from about 5 to 90+ minutes; the cursor is now iterated in itersize batches, cutting a 100,000-row export from 100,001 fetches to 51, checked against Google's GCP system test before merge and closing issue #72075 (5ec2045).
google/adk-python #6419 · #6705 ✅ shipped in Google releases
A Windows-path fix and an eval regression testadk eval read the colon in C:\ as a case selector, so Windows paths broke; fixed the eval-case parser (6f6106f, shipped in v2.6.0). Eval tools already received EvalCase.session_input.state, but instruction templates such as {some_key} historically did not during evaluation; added a LocalEvalService regression test so that path cannot silently break again, closing issue #5037 (3983fa1, shipped in v2.10.0). Google lands patches through Copybara, so both PRs read closed while the commits sit on main, authored by me.
apache/magpie #1089 ✅ merged
fix(license-compliance-audit): handle large blobs and canonical SPDXFiles over ~1 MiB return without inline content, so the audit silently mis-flagged them as non-compliant. Fetches via raw-media API, separates uninspected from violating. Closed issue #944.
numpy/numpy #32386 ✅ merged
BUG: close duplicated file descriptor if fdopen failsC-level fix for a descriptor leak in NumPy's error path, with regression tests handling WASM/musl. All 91 required CI checks passed.
topoteretes/cognee #4126 ✅ merged
fix(cli): reject dry runs in API dispatch mode--dry-run was silently ignored when --api-url was set, executing real remote operations.
corsairdev/corsair #1200 ✅ merged
feat(pinecone): production-grade integration48 operations across four API surfaces, typed Zod schemas, dynamic host routing. Landed as f7820d6, +3,948 across 26 files, closing issue #1199.
risa-labs-inc/BossConsole #324 · #325 · #477 · #508 · #509 ✅ merged
Five changes in a Kotlin desktop appPlugin class loading made parallel-safe by serializing same-name first loads while independent classes still load concurrently (56de814); unloaded plugins stopped from resolving resources through the host (895c57e); heavyweight overlays made to honor RTL layouts and round fractional sizes instead of truncating, removing one-pixel drift at fractional display scales (21cc371); a newer screen-capture request now cancels the callback it supersedes, so stale picker actions can no longer clear or resolve the request on screen (beca66c); and repository-wide convention scans hardened so generated build/ trees cannot affect their results (89a6137).

Systems & software

conveyor-job-queue: a Go job queue that lost no jobs across 50 kill -9 trials

Repo · Go Connect RPC Protobuf Prometheus

crash-safety throughput

Durable distributed job queue built around crash-safety: a CRC32C-checksummed, length-prefixed write-ahead log with torn-write detection, flushed with F_FULLFSYNC rather than plain fsync; group commit so N concurrent submitters share one flush; lease and fencing logic with jittered backoff and dead-lettering. 32,232 submissions/sec, 7.7 ms recovery for 100K records.

The hard part: a worker that stalls for 30 seconds and comes back must not be able to acknowledge a job someone else now owns. The badges above are regenerated by CI, not typed here.

redoscope: a ReDoS analyzer with 0 false positives and 0 misses on 627 real regexes

Repo · TypeScript Node ≥22.18 zero deps

ReDoS static analysis that verifies its own findings. Parses the full ECMAScript regex grammar, compiles to an NFA modelling real backtracking, detects exponential and polynomial paths via product automata, then generates a witness attack string and times it inside a killable child process. Benchmarked on 627 real-world regexes: a star-height heuristic gave 10 false positives and missed 14 real vulnerabilities; this pipeline hit 0 false positives and 0 misses.

The hard part: "looks risky" is not a vulnerability. The tool has to produce a concrete string that actually hangs the engine, and prove it with a measured growth curve.

forgegraph: a Go build engine where one edit reruns 4 of 11 tasks and a warm rebuild takes 1 ms

Repo · Go React Flow SSE

Local incremental build engine in a single binary. Validates the task DAG with cycle paths, runs independent tasks on a bounded parallel scheduler, and keys each cache entry by SHA-256 over command, environment digest, input contents and dependency fingerprints — content, not mtime. Every miss names the input that changed, and an embedded dashboard streams the live graph, concurrency timeline and cache decisions.

walscope: a SQLite WAL inspector that never opens the database through SQLite

Repo · Python SQLite internals

Read-only inspector for SQLite write-ahead-log state. Parses the database header, -wal frames (salts, rolling checksums, commit markers) and the -shm wal-index as plain files to report mxFrame, nBackfill, checkpoint lag and read marks. A built-in lab reproduces a pinned reader stalling the checkpoint (lag 25 frames) and its release (lag 0).

The hard part: it never calls sqlite3.connect() on the target — even a read-only open can run recovery and rewrite the very files being inspected. Tests enforce it.

mutant: browser mutation testing that caught 11 of 17 bugs a fully passing suite missed

Repo · Live demo · TypeScript acorn Web Workers

A hand-written AST engine injects real bugs — flipped operators, nudged boundaries, wiped constants, dropped guards — at the exact token, re-parses every mutant to guarantee valid JavaScript, and runs each against your tests in a disposable Web Worker with a hard timeout. Survivors come back as one-line diffs and every kill is attributed to the test that caught it. On the bundled GST example, a fully passing suite scores 35%.

Teachable Voice: an Android assistant taught by one demonstration, 18/18 commands understood

Repo · Kotlin Android Accessibility

Say a command, perform it once in any app, and it replays on voice — with new values ("a Farmhouse instead", "two of them", "deliver to Work"), new wording and Hinglish. Runs only through the Accessibility Service, no app APIs or deep links; a tap-capture overlay records Compose and web-view taps that accessibility events miss. Replay is scored element matching with no network call, the language model steps in only when the screen differs, and a payment/OTP/login guard runs on every action. 7/7 cross-app runs (taught on Amazon, run on Myntra), 39 of 40 replay steps without the model.

shramshield: a heat-safety shift planner built on an ISO 7243 WBGT engine, 83 tests

Repo · Live · TypeScript React Vitest

Turns a weather forecast into an enforceable heat-safety shift plan for outdoor crews: shift window, work/rest minutes per hour, water, and stop-work hours. A framework-free WBGT engine implements ISO 7243 and ACGIH screening limits and an optimiser ranks candidate windows — in the demonstrated Delhi forecast, moving the shift to 05:00 cuts time above the safe limit from 240 to 60 minutes.

rescuerelay: crisis-resource matching that explains every recommendation

Repo · Next.js 15 TypeScript Leaflet Recharts

Camps and clinics report needs, providers register available resources, and a deterministic matching engine ranks compatible pairings by priority, distance, availability and fairness — with a plain-language explanation attached to every recommendation. Offline-capable field intake for low-connectivity reporting.

LiveEnv (private): a local-first geospatial platform with a double-entry ledger

Next.js SQLite WebSocket IndexedDB MapLibre

Decay-based geospatial social platform across four surfaces. Geohash-partitioned SQLite serving a hot TTL-swept table alongside durable transactional tables; paid placements settled through an append-only double-entry ledger so money never moves without minting the pin; IndexedDB offline outbox with idempotent client-ID-keyed sync.

AgentBar (private): a native macOS menu bar app tracking six AI coding agents

Swift Xcode Keychain SQLite

Unifies usage tracking across Codex, Claude Code, Gemini, Cursor, OpenCode and Z.ai into one stacked bar with per-service metrics. Secure credential handling through macOS Keychain across heterogeneous sources; shipped as a signed .dmg under MIT.

How Conveyor survives a kill -9

stateDiagram-v2
    direction LR
    [*] --> Pending: Submit, appended to WAL and flushed before ack
    Pending --> Leased: Lease granted, epoch incremented
    Leased --> Done: Ack carrying the current epoch
    Leased --> RetryWait: Nack, or lease timeout reclaim
    RetryWait --> Pending: backoff with jitter
    RetryWait --> DeadLetter: retry budget exhausted
    Leased --> Leased: Ack carrying a stale epoch — rejected
    Done --> [*]
    DeadLetter --> [*]
Loading

The epoch is the fencing token. A lease timeout and an explicit Nack decrement the same retry budget, so a poison-pill job cannot loop forever without reaching the dead-letter queue.


AI engineering

Netra: an offline screen assistant for blind users on the Snapdragon NPU, in English and Hindi

Repo · Python ONNX Runtime QNN Whisper VLM

Press a key and ask about the screen, a scanned bill, or a medicine strip held to the webcam: Whisper on the NPU transcribes, a vision-language model reasons over the screenshot with OCR grounding, and the answer is spoken sentence by sentence while it is still generating, interruptible at any point. Nothing leaves the laptop — which is the point on banking, Aadhaar and medical screens that cloud describers would upload.

BlindSpot: a fraud-decline auditor that replays a 590,540-row benchmark byte-for-byte

Repo · Python scikit-learn Streamlit pytest

Uncertainty-aware auditor for fraud-model declines under censored outcomes. Freezes a verification policy before outcomes are revealed, estimates false-decline rates with known sampling propensities, and reports when the review budget is too small to support a stable claim. 40 tests keep sealed labels out of the product path.

WebMCP Incident Workspace: an agent incident room where irreversible actions stay human-only

Repo · Live demo · WebMCP JavaScript Vercel

Human-agent incident command room with nine browser-native tools. Agents can compare deterministic remediation futures, but stale plans are rejected and irreversible actions remain human-only; successful execution emits a retrievable decision receipt. Verified through seven real-client scenarios and 18 automated tests.

modelgauntlet: SHIP / FIX / BLOCK release verdicts where AI never grades AI

Repo · Next.js Zod Ajv Vitest

Pre-release testing platform that stress-tests structured AI tasks across open-source models and returns a deterministic verdict. One rule enforced end to end: AI proposes, deterministic TypeScript code decides. 23 automated tests across parsing, assertion, schema-validation and verdict boundaries.

multimodal-evidence-review: claim accuracy lifted from 70% to 85%, top 0.25% of 15,000+ at HackerRank Orchestrate

Repo · Python Gemini 2.5 Flash

Multimodal damage-claim adjudication pipeline — model output validated and repaired against a strict 14-column schema with tightly constrained enums. Semantic consistency rules lifted claim-status accuracy with zero additional model calls. SHA-256 content-addressed caching, resumable batch pipeline, graceful stop on quota errors. Top 0.25% of 15,000+ registrants at HackerRank Orchestrate, June 2026.

Synapse Memory Lab: DataForge 2026 finalist at IIT Kharagpur, an interactive essay on BDH-GPU memory

Repo · Live · TypeScript Canvas

Three live labs show causal linear attention computed as Hebbian outer-product writes into a fixed-shape synaptic state, check that recurrent form numerically against an explicit-history oracle, and push an associative memory past capacity so interference becomes visible.

hospital-readmission-risk: a top-20% call list that catches 40% of readmissions

Repo · Python XGBoost LightGBM SHAP FastAPI

30-day readmission model framed as a capacity question — if the team can call 20% of today's discharges, which 20%? Calling the top-risk fifth catches 40.2% of readmissions (2.05× lift). ROC-AUC 0.684 with patient-clustered bootstrap intervals, a seven-split robustness check that reports the published split as the most favourable, decision-curve analysis, SHAP reason codes, a scoring API and 27 tests.

kaamtwin: a scheduling twin proven by 15 anti-hardcoding causal tests

Repo · Next.js TypeScript

Causal digital-twin simulator that reorders manufacturing order queues by deadline priority — 62 → 59 late deliveries in the demonstrated scenario, verified by causal tests (reversed-input, capacity, duration, immutable-hash) that rule out a hardcoded result. 67 automated tests at 88%+ coverage, 8/8 Chromium e2e, hardened CSP.

zero-token-router: 19/19 evaluations passed inside a 4 GB, 2-vCPU image

Repo · Python Fireworks AI API

Hybrid router resolving deterministic tasks locally and escalating only genuinely hard work to Fireworks-hosted models, shipped as a public Linux/amd64 image.


Experience

AI Intern — Bharat Electronics Limited, Bengaluru · Jul – Aug 2026
Backend API integrations in FastAPI, a React/Vite analytics dashboard turning application data into usage, latency and performance views, and improvements to the project's RAG pipeline.

Remote AI Intern — Facillima · May – Jun 2026
Shipped an AI-native EdTech platform end to end in Next.js/TypeScript — technical SEO, structured data, analytics instrumentation, third-party LMS integration — and led evaluation of the company's LLM observability stack.

Certifications — Anthropic AI Fluency, MCP, Agent Skills and Subagents · Walmart Global Tech Advanced Software Engineering · Deloitte Cyber · Google Ads Apps


Toolkit

AI & agents: runtimes, protocols and evaluation tooling

Concepts: RAG · autonomous agents · multi-agent orchestration · deterministic verification · model evaluation & benchmarking · release gating · on-device NPU inference · prompt engineering · LLM observability

Tools: LangSmith · Braintrust · Zod · Ajv · Vitest · agent skills & persistent memory

Machine learning & data: modelling, explainability and evaluation

Concepts: ensemble methods · CNNs & Grad-CAM explainability · feature engineering · model calibration · ROC-AUC / permutation importance · SHAP · decision-curve analysis

Systems & infrastructure: durability, scheduling and delivery

Concepts: write-ahead logging · crash recovery & durability · fencing tokens · lease-based dispatch · group commit · content-addressed caching · incremental builds · static analysis / automata · mutation testing · local-first & offline sync

Backend, frontend & stores: services, interfaces and databases

Core CS: data structures & algorithms · DBMS · object-oriented programming · distributed systems · software testing (pytest, Vitest, Playwright, Ruff)


Conveyor's crash-safety and throughput figures are re-measured on every push by verify. The header and project cards are generated by the scripts in tools/.

Pinned Loading

  1. conveyor-job-queue conveyor-job-queue Public

    Durable distributed job queue in Go: crash-safe WAL, at-least-once delivery with fencing tokens, lease-timeout reclaim, retries with jittered backoff, and a dead-letter queue.

    Go

  2. redoscope redoscope Public

    ReDoS static analysis that proves its own findings — compiles regexes to an NFA, detects exponential paths via product automata, then generates and times a witness attack. 0 false positives and 0 m…

    TypeScript

  3. kaamtwin kaamtwin Public

    Evidence-linked deterministic workflow debugger for informal microbusiness operations

    TypeScript

  4. modelgauntlet modelgauntlet Public

    Break your AI before users do — deterministic release evidence across open models.

    TypeScript

  5. multimodal-evidence-review-orchestrate multimodal-evidence-review-orchestrate Public

    Multimodal damage-claim adjudication pipeline — model output validated and repaired against a strict 14-column schema; semantic consistency rules lifted claim-status accuracy from 70% to 85% with z…

    Python

  6. zero-token-router zero-token-router Public

    Hybrid LLM router that resolves deterministic tasks locally and escalates only genuinely hard work — 19/19 simulated evaluations passed in a public Linux/amd64 image under 4 GB RAM and 2 vCPU.

    Python