#attention
7 posts tagged
[AI Search] QKV Projection Sharing: Transformer는 정말 Q, K, V 세 개가 모두 필요할까
Do Transformers Need Three Projections? 논문을 AI Search stack 관점에서 읽는다. Query, Key, Value projection을 일부 공유해도 품질을 크게 잃지 않으면서 KV cache를 줄일 수 있는지, Q-K=V가 왜 실용적인 절충점인지, GQA/MQA와 결합하면 on-device·edge search serving에 어떤 의미가 생기는지 정리한다.
AI 엔지니어 필독 논문 10개 — ① 기초 아키텍처 (Attention, VAE, GANs)
AI 면접 단골 논문 10개를 정리한 시리즈 첫 편. Attention Is All You Need, VAE, GANs가 어떻게 현대 AI의 기초를 다졌는지 이해한다.
[논문 리뷰] DeepSeek-V4 — 1M Context에서 KV Cache 10% 수준으로 압축한 Hybrid Attention
DeepSeek-V4는 기존 Multi-Head Attention의 개념을 바탕으로, Compressed Sparse Attention(CSA)와 Heavily Compressed Attention(HCA)을 결합한 Hybrid Attention으로 1M token context를 지원하면서도 KV cache를 90% 감축했다. 이전 Attention 이해하기 시리즈의 Q/K/V와 Multi-Head Attention 개념을 이어 DeepSeek-V4가 어떻게 구현했는지 살펴본다.
[NLP] Transformer 3가지 Attention 자세히 보기 (Encoder/Decoder Self-Attention, Cross-Attention, Multi-Head)
이전 글에서 등장한 Transformer의 3가지 Attention(Encoder Self-Attention, Decoder Masked Self-Attention, Encoder-Decoder Attention)이 각각 어떻게 동작하는지, 그리고 Multi-Head Attention이 왜 필요한지 정리한다.