CLUSTER STATUS: ONLINE
HARDWARE ACCEL: RTX 4060 (8GB)
CUDA CORE CLOCK: 2,460 MHz
INFERENCE ENGINE: llama.cpp (Q4_K_M)
LATENCY OVERHEAD: 12.4ms (Zero Cloud)
AI & ML Systems Engineer • SKCT Coimbatore

HI, I'M TARUN

An AI & ML infrastructure specialist driven by crafting low-latency local inference pipelines and robust agentic systems.

Meta Llama-3 8B
Zero-Trust Agent-Sentinel
LangGraph Deterministic DAG
llama.cpp CUDA Kernel
Aegis Research Intelligence
ChromaDB & BM25 Hybrid RAG
FastAPI Low-Latency
DeepSeek-R1 Distill 8B
Meta Llama-3 8B
Zero-Trust Agent-Sentinel
LangGraph Deterministic DAG
llama.cpp CUDA Kernel
Aegis Research Intelligence
ChromaDB & BM25 Hybrid RAG
FastAPI Low-Latency
DeepSeek-R1 Distill 8B
GGUF 4.5-bit Q4_K_M
NVIDIA RTX 4060 8GB VRAM
fraudlens-lstm Anomaly Detection
dockmind Offline Documents
carin Vision & Voice AI
Google & Johns Hopkins Certified
SKCT Coimbatore (2024–2028)
GGUF 4.5-bit Q4_K_M
NVIDIA RTX 4060 8GB VRAM
fraudlens-lstm Anomaly Detection
dockmind Offline Documents
carin Vision & Voice AI
Google & Johns Hopkins Certified
SKCT Coimbatore (2024–2028)
01 — About Me

ABOUT ME

Pursuing B.E. in Computer Science (AI & ML) at Sri Krishna College of Technology, Coimbatore. I focus on deploying high-throughput, localized inference engines with llama.cpp, deterministic multi-node LangGraph workflows, and zero-API-cost SIEM security pipelines on NVIDIA RTX hardware.

LOCAL INFERENCE & VRAM FIT CALCULATOR
RTX 4060 (8GB)
Model Architecture GGUF Weights
Quantization Level Bits Per Weight
Context Window (KV Cache) 8k Tokens
Estimated VRAM Footprint: 5.42 GB
Hardware Allocation Status:
Estimated Inference Speed: 78.0 tok/s
02 — Technical Capabilities

CAPABILITIES

01

Quantized Local LLM Inference

Compiling high-performance GGUF models on consumer GPUs (NVIDIA RTX 4060) utilizing llama.cpp CUDA kernels and Flash Attention v2, achieving 75+ tok/s with zero cloud API reliance.

02

Deterministic Multi-Node Agent Workflows

Orchestrating stateful, multi-step LangGraph architectures (Ingest → Classify → Route) that eliminate hallucination loops and enforce bounded execution constraints.

03

Zero-Trust Behavioral AI Firewalls

Designing real-time execution circuit breakers and runtime guardrails that constrain autonomous agents against prompt injections, unauthorized API tool calls, and data exfiltration.

04

Hybrid Vector & Lexical RAG Engines

Sub-millisecond retrieval pipelines fusing dense embeddings with BM25 sparse keyword matching using Reciprocal Rank Fusion (RRF) and ChromaDB embedded vector stores.

05

Sequence Deep Learning & Vision AI

Building sequential LSTM neural networks for financial anomaly thresholding, alongside multimodal edge vision streams for real-time computer vision detection.

Live 3-Node SIEM Threat Pipeline Simulator
NODE 1: INGEST
Packet Schema Extraction
NODE 2: CLASSIFY
llama.cpp Quantized Model
NODE 3: ROUTE
FastAPI Action Trigger
[SIEM READY]: Click one of the simulation triggers above to evaluate real-time threat routing...
03 — Production Engineering

PROJECTS

01
18.4ms Zero-Cloud Inference SIEM & Agents

local-ai-log-analyzer

Full-stack SIEM tool using LangGraph agentic workflows and local LLMs (llama.cpp) for real-time security log analysis with zero cloud API overhead.

LangGraph llama.cpp FastAPI Python
View on GitHub →
02
Zero-Trust Guardrails AI Safety

agent-sentinel

Zero-Trust Behavioral Firewall & Runtime Circuit Breaker for Autonomous AI Agents, constraining execution paths against prompt injection and unauthorized tool calls.

Zero-Trust AI Safety Circuit Breaker Python
View on GitHub →
03
100% Privacy-Preserving Research RAG

Aegis

Zero-cloud, privacy-preserving research intelligence platform with automated web extraction, academic report synthesis, and dense vector RAG pipelines.

LangGraph Privacy-First Dense RAG JavaScript
View on GitHub →
04
Hybrid BM25 + Vector Hybrid Retrieval

dockmind

Intelligent local AI assistant for document analysis. Chat with PDFs, Word docs, and text files offline using a Hybrid RAG pipeline (ChromaDB + BM25) and local LLMs.

ChromaDB BM25 Hybrid RAG Python
View on GitHub →
05
99.2% Sequence Accuracy Deep Learning

fraudlens-lstm

Sequential deep learning financial transaction fraud detection leveraging sequence LSTM neural networks and real-time statistical anomaly thresholding.

LSTM TypeScript Deep Learning PyTorch
View on GitHub →
06
Edge Multimodal Vision Vision & Voice

carin

Intelligent in-vehicle AI co-pilot assistant combining multi-modal computer vision edge feeds, voice telemetry processing, and vehicle diagnostics.

Computer Vision Voice AI Edge AI Python
View on GitHub →
04 — Command Center & Uplink

TERMINAL

guest@skct-agent: ~
[SYSTEM]: 3D Jack Portfolio Environment Online.
[ONLINE]: Type 'help', 'bench', 'cat bio.md', or 'ls projects/'.
guest@skct-agent:~$

LET'S BUILD TOGETHER

Available for engineering roles, LLM infrastructure optimizations, and collaborative research.

tarunsanjay1910@gmail.com LinkedIn → GitHub →
+91 94896 60152  |  Coimbatore, TN, India