İçeriğe atla
0
  • Ders Ara
  • Ana Sayfa
  • Kategoriler
    • All Categories
      • Individual Categories
    • Gruplar
    • Okunmamış 0
    • Güncel
    • Kullanıcılar
    • Hakkımızda
    • Öğrenci Fırsatları
    • Akademik Takvim
    • CV oluşturucu
    • IEU Timetable
    • Devamsızlık App
    • IEU GPA Hesaplayıcı
    • Niki Cüzdan
    • Ders Ara
    • Ana Sayfa
    • Kategoriler
      • All Categories
        • Individual Categories
      • Gruplar
      • 0 Okunmamış 0
      • Güncel
      • Kullanıcılar
      • Hakkımızda
      • Öğrenci Fırsatları
      • Akademik Takvim
      • CV oluşturucu
      • IEU Timetable
      • Devamsızlık App
      • IEU GPA Hesaplayıcı
      • Niki Cüzdan
      Daralt
      IEU Forum – İzmir Ekonomi Üniversitesi Öğrenci Topluluğu Platformu

      IEU Forum

      -- çevrimiçi
      1. Ana Sayfa
      2. Large Language Models
      3. paper: ZEDA for MoE

      Final Unicourse'tan Çalış, Yüksek Notu Garantile!

      %25 İndirim Kodu: FRM25
      Yükleniyor...
      Dersi İzle
      GÖRÜNTÜLEYENLER
      +36
      Premium Özellik
      Bu konuyu kimlerin görüntülediğini görmek için Premium üyelik gerekir.
      Premium'a Geç

      Vizesine Unicourse'tan Çalış, Yüksek Notu Garantile!

      A B C D Çıkmış Sorular Formül Kağıtları Konu Anlatımı Sınav İpuçları Örnek Sınav
      Dersi İzle
      YENİ ÖZELLİK

      Bi'Öğrenci Fırsatları
      Forum'da!

      Bi'Öğrenci ile artık forum üzerinden en güncel indirimlere, anlık fırsatlara ve avantajlı tekliflere ulaşabilirsin.

      FIRSATLARI KEŞFET
      Red Bull Basement
      SPONSORLU ETKİNLİK

      Fikrini Gerçeğe Dönüştür

      Projeni dünyaya göstermek için sahne hazır. Red Bull Basement başvuruları açık.

      Başvurunu Yap

      🎉 Foruma Yeni Özellik Geldi!

      Sizin için PDF toollarını getirdik!

      İncele ve Kullan

      paper: ZEDA for MoE

      Konu Zamanlandı Sabitlendi Kilitli Taşındı Large Language Models
      llm
      1 İleti 1 Yayımlayıcılar 0 Bakış
      • En eskiden en yeniye
      • En yeniden en eskiye
      • En çok oylanan
        Cevap
        • Yeni başlık oluşturarak cevapla
        Cevaplamak için giriş yapın
        Bu başlık silindi. Sadece başlık düzenleme yetkisi olan kullanıcılar görebilir.
        • leanleft@lemmy.mlL This user is from outside of this forum
          leanleft@lemmy.mlL This user is from outside of this forum
          leanleft@lemmy.ml
          yazdı Son düzenleyen:
          #1

          arxiv Post-Trained MoE Can Skip Half Experts via Self-Distillation

          https://huggingface.co/TsinghuaC3I Tsinghua University

          Mixture‑of‑Experts (MoE) basics
          An MoE layer contains many “expert” feed‑forward sub‑networks, but for each token only a small subset is activated by a router. This sparse activation lets the overall model grow very large while keeping the per‑token compute bounded. Dynamic MoE variants go further by letting the router decide, input‑dependently, how many experts to use, reducing computation for easy tokens.

          Problem the paper addresses
          Existing dynamic‑MoE techniques usually require training the model from scratch or performing task‑specific fine‑tuning. Consequently, a fully trained static MoE model cannot be easily turned into a dynamic one without risking loss of the routing knowledge that was already learned. This limits practical deployment because inference costs remain high even when many tokens could be handled with fewer experts. <citation src="1"></citation>

          How the new method (ZEDA) works

          1. Zero‑Expert injection – a parameter‑free “zero‑output” expert is added to every MoE layer. These experts produce no contribution unless selected, allowing the model to skip computation for certain tokens.
          2. Two‑stage self‑distillation –
            • Stage 1 (SFT): the augmented model is fine‑tuned while staying close to the original outputs.
            • Stage 2 (OPD): the original static MoE acts as a frozen teacher; the student (augmented model) learns via self‑distillation, guided by a group‑level balancing loss that keeps expert loads even.
              This stabilises the conversion from static to dynamic architecture without needing a new pre‑training run. <citation src="2"></citation>

          Benefits emphasized by the authors

          • Computation reduction: > 50 % of expert FLOPs are eliminated, meaning many tokens bypass expert computation entirely.
          • Small accuracy loss: performance drops only marginally across a suite of 11 benchmarks (math, code, instruction following).
          • Speed‑up: about 1.2 × end‑to‑end inference acceleration is observed.
          • Low adaptation cost: ZEDA requires only a brief self‑distillation fine‑tuning phase, avoiding the expense of full re‑training.
          • Compatibility: because zero‑experts are parameter‑free, the method can be applied to existing large MoE models without architectural redesign. <citation src="4"></citation>
          1 Cevap Son cevap
          1

          Hello! It looks like you're interested in this conversation, but you don't have an account yet.

          Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

          With your input, this post could be even better 💗

          Kayıt Ol Giriş
          Cevap
          • Yeni başlık oluşturarak cevapla
          Cevaplamak için giriş yapın
          • En eskiden en yeniye
          • En yeniden en eskiye
          • En çok oylanan


            Önerilen Başlıklar

            • O

              Microsoft Paint embeds hidden IDs in locally generated AI images

              Takip ediliyor Susturulmuş Konu Zamanlandı Sabitlendi Kilitli Taşındı LLM llm
              1
              1
              1 Oy
              1 İleti
              0 Bakış
              Kimse yanıtlamadı
            • O

              The Invisible Font Resume Hack is Breaking AI Job Screeners. As a Recruiter, I Don’t Blame Candidates

              Takip ediliyor Susturulmuş Konu Zamanlandı Sabitlendi Kilitli Taşındı LLM llm
              1
              1 Oy
              1 İleti
              0 Bakış
              Kimse yanıtlamadı
            • O

              I lost my grandmother

              Takip ediliyor Susturulmuş Konu Zamanlandı Sabitlendi Kilitli Taşındı LLM llm
              1
              1
              1 Oy
              1 İleti
              0 Bakış
              Kimse yanıtlamadı
            • Q

              New Deepseek V4 Pro model has been released

              Takip ediliyor Susturulmuş Konu Zamanlandı Sabitlendi Kilitli Taşındı Large Language Models llm
              1
              1
              1 Oy
              1 İleti
              0 Bakış
              Kimse yanıtlamadı
            • quokka@quokk.auQ

              Residents call for moratorium on AI data centres at meeting in Darwin

              Takip ediliyor Susturulmuş Konu Zamanlandı Sabitlendi Kilitli Taşındı Australia ai australia llm aus canberra
              1
              1 Oy
              1 İleti
              3 Bakış
              Kimse yanıtlamadı

            Developed by Enes Uysal & Kadir Ay

            4

            Çevrimiçi

            8.8k

            Kullanıcı

            1.9k

            Konu

            3.7k

            İleti
            • Giriş

            • Hesabınız yok mu? Kayıt Ol

            • Aramak için giriş yapın veya kaydolun
            • İlk ileti
              Son ileti