Back to "Κατασκευή ενός "δικηγόρου GPT" για το blog σας - Μέρος 6: Τοπική ενσωμάτωση LLM"

This is a viewer only at the moment see the article on how this works.

To update the preview hit Ctrl-Alt-R (or ⌘-Alt-R on Mac) or Enter to refresh. The Save icon lets you save the markdown file to disk

This is a preview from the server running through my markdig pipeline

AI AI-Article C# GGUF LLamaSharp LLM mostlylucid.blogllm

Κατασκευή ενός "δικηγόρου GPT" για το blog σας - Μέρος 6: Τοπική ενσωμάτωση LLM

Wednesday, 12 November 2025

ΠΡΟΕΙ∆ΟΠΟΙΗΣΗ: ΤΙΣ ΠΡΟΘΕΣΜΙΕΣ ΠΟΥ "ΠΕΡΙΛΑΜΒΑΝΟΝΤΑΙ."

Πιθανότατα πολλά από αυτά που είναι κάτω δεν θα δουλέψουν; Εγώ τα δημιουργώ ως το πώς-για μένα και στη συνέχεια να κάνουμε όλα τα βήματα και να πάρει το δείγμα εφαρμογή εργασίας...Έχετε sneaky και να τους δει! Θα είναι πιθανώς έτοιμα μέσα Δεκεμβρίου.

## Εισαγωγή

Καλώς ήρθατε στο 6ο μέροςΚατασκευάσαμε την πλήρη υποδομή - αγωγό κατάποσης (Μέρος 4), Windows client (Μέρος 5), ε p i ι χ ε ι ρ ή σε ι ς και p i ρ α γ μ α το p i οι ή σε ι ς (Μέρος 3), και ρύθμιση GPU (Μέρος 2

Τώρα έρχεται το συναρπαστικό μέρος: ενσωμάτωση ενός τοπικού LLM για να δημιουργήσει πραγματικά προτάσεις γραφής.

ΣΗΜΕΙΩΣΗ: Αυτό είναι μέρος των πειραμάτων μου με AI (βοηθητική σύνταξη) + δικό μου μοντάζ.

Ίδια φωνή, ίδιος πραγματισμός, απλά γρηγορότερα δάχτυλα.

Εδώ είναι που επιτέλους κάνουμε το "AI" μέρος του "AI βοηθός γραφής" εργασία.

Θα τρέξουμε μεγάλα γλωσσικά μοντέλα τοπικά σε A4000 GPU σας, δημιουργώντας προτάσεις context-aware με βάση τις προηγούμενες δημοσιεύσεις blog σας. |--------|-----------|-------------------| | Γιατί το τοπικό LLM; | ✅ Complete | ❌ Data sent to third party | | Πριν καταδυθούμε, ας καταλάβουμε γιατί τρέχουμε τοπικά μοντέλα αντί να χρησιμοποιούμε το API του OpenAI. | ✅ Free after setup | ❌ Per-token pricing | | Τοπική σύγκριση εναντίον API | ✅ <1 second | ⚠️ Network dependent | | Τοπικό LLM API (OpenAI κ.λπ.) | ✅ Full control | ❌ Limited | | Απόρρητο | ✅ Any GGUF model | ❌ Provider's models only | | Κόστος | ✅ Works offline | ❌ Requires internet | | Ευελιξία | ❌ Complex | ✅ Simple |

Προσαρμογή

Επιλογή μοντέλου

Εκτός σύνδεσηςName

graph TB
    A[C# Application] --> B{Integration Method}

    B --> C[LLamaSharp]
    B --> D[ONNX Runtime]
    B --> E[TorchSharp]
    B --> F[HTTP API]

    C --> G[llama.cpp bindings]
    G --> H[GGUF Models]

    D --> I[ONNX Models]
    I --> J[Limited Model Support]

    E --> K[PyTorch Models]
    K --> L[Complex Setup]

    F --> M[External Process]
    M --> N[Ollama, LM Studio]

    class C recommended
    class G,H llamaSharp

    classDef recommended stroke:#333,stroke-width:4px
    classDef llamaSharp stroke:#333,stroke-width:2px

Ρύθμιση

Για βοηθό γραφής, ιδιωτικότητα και θέμα κόστους.

  • Δεν θέλουμε τα σχέδια blog που αποστέλλονται σε εξωτερικούς APIs, και per-token τιμολόγηση προσθέτει γρήγορα για ένα καθημερινό εργαλείο γραφής.
  • Επιλογές ενσωμάτωσης LLM για C#
  • Υπάρχουν διάφοροι τρόποι για να τρέξει LLMs σε C#:
  • Η επιλογή μου: LLamaSharp
  • Γιατί;

Ιθαγενείς C# δεσίματα για llama.ccpp (γρήγορη βιβλιοθήκη συμπερασμάτων)

Υποστηρίζει μορφή GGUF (σύγχρονα, ποσοτικοποιημένα μοντέλα)

Ενσωματωμένη επιτάχυνση CUDAΕνεργή ανάπτυξη και μεγάλη κοινότητα

graph LR
    A[Original Model<br/>Llama 2 7B<br/>~28GB float32] --> B[Quantization]

    B --> C[Q4_K_M<br/>~4.1GB<br/>4-bit]
    B --> D[Q5_K_M<br/>~4.8GB<br/>5-bit]
    B --> E[Q6_K<br/>~5.5GB<br/>6-bit]
    B --> F[Q8_0<br/>~7.2GB<br/>8-bit]

    C --> G[Fast, Lower Quality]
    D --> H[Balanced]
    E --> I[Higher Quality]
    F --> J[Near Original]

    class A original
    class C,D quantized
    class H recommended

    classDef original stroke:#333,stroke-width:2px
    classDef quantized stroke:#333,stroke-width:2px
    classDef recommended stroke:#333,stroke-width:2px

Δουλεύει με Llama, Mistral, Phi, Gemma, και πολλά άλλα:

  • Κατανόηση πρότυπων μορφοτύπων & κβαντισμού
  • Μορφή GGUF
  • GGUF
  • (GPT-Generated Unified Format) είναι το πρότυπο για την αποτελεσματική λειτουργία LLMs.

Εξήγησε η κοστολόγηση

Original model: 32-bit πλωτά (πολύ μεγάλα, πολύ ακριβή) |-------|---------------|------------|-----------|------------|------------|---------| | Q4: 4-bit ακέραιοι (75% μικρότεροι, ελάχιστη απώλεια ποιότητας) | 2.3GB | ~4GB | ✅ Easy | ✅ Easy | ✅ Easy | ⭐⭐⭐ Good | | Q5/Q6: Γλυκό σημείο για τις περισσότερες περιπτώσεις χρήσης | 4.1GB | ~6GB | ✅ Tight | ✅ Good | ✅ Easy | ⭐⭐⭐ Good | | Q8: Κοντά αρχική ποιότητα, ακόμα 4x μικρότερη | 4.1GB | ~6GB | ✅ Tight | ✅ Good | ✅ Easy | ⭐⭐⭐⭐ Better | | Model Selection by Hardware | 4.1GB | ~6GB | ✅ Tight | ✅ Good | ✅ Easy | ⭐⭐⭐⭐ Better | | • Μοντέλο Μέγεθος (Q4_K_M) • Χρήση VRAM Ταιριάζει με 8GB; • Ταιριάζει με 12GB; • Ταιριάζει με 16GB; • Ποιότητα | 4.7GB | ~7GB | ⚠️ Very Tight | ✅ Good | ✅ Easy | ⭐⭐⭐⭐⭐ Best | | Phi-3 Mini (3.8B) | 7.4GB | ~10GB | ❌ No | ⚠️ Tight | ✅ Good | ⭐⭐⭐⭐ Better |

Llama 2 7B

  • Mistral 7BGemma 7BLlama 3 8BLlama 2 13B**Συστάσεις της GPU:**8GB VRAM
  • : Ξεκινήστε με: Mistral 7BήPhi-3 Mini(ασφαλέστερος)
  • 12GB VRAM: Llama 3 8B(βέλτιστη ποιότητα) ήMistral 7B
  • **(γρήγορο)**16GB VRAM (το στήσιμό μου)

Llama 3 8B: ή προσπαθήστεΜοντέλα 13ΒΜόνο CPU: Κάθε μοντέλο λειτουργεί, απλά πολύ πιο αργά (ξεκινήστε με Phi-3 Mini για ταχύτητα)

  • Η πρότασή μου
  • Mistral 7B
  • (τελευταία έκδοση) ή
  • Λάμα 3

8B

Εξαιρετική ποιότητα για την τεχνική γραφή

Λειτουργεί σε όλα τα μεγέθη GPU

  1. Αρκετά γρήγορο για διαδραστική χρήσηΚαλό να ακολουθείς τις οδηγίες
  2. Downloading Models"mistral 7b gguf"
  3. Τα μοντέλα διανέμονται στο Hugging Face.

**Θα χρησιμοποιήσουμε ποσοτικοποιημένες εκδόσεις GGUF.**Βρίσκοντας μοντέλα GGUF

Ψάξτε για τις ποσοτικοποιήσεις του TheBloke (πιο δημοφιλής)

# Install huggingface-cli
pip install huggingface-hub

# Download Mistral 7B Q5_K_M (recommended)
huggingface-cli download TheBloke/Mistral-7B-Instruct-v0.2-GGUF \
    mistral-7b-instruct-v0.2.Q5_K_M.gguf \
    --local-dir C:\models\mistral-7b \
    --local-dir-use-symlinks False

Απευθείας σύνδεσμοι

  1. (Οι ποσοτικοποιήσεις του TheBloke):
  2. Mistral-7B-Instruct-v0.2-GUFmistral-7b-instruct-v0.2.Q5_K_M.ggufLlama-2-7B-Chat-GGUF
  3. Llama-3-8B-Instruct-GUF
  4. Κατεβάστε το ειδικό QuantizationC:\models\mistral-7b\

Ή κατεβάστε με το χέρι:

Κάντε κλικ στην καρτέλα "Φάκελοι και εκδόσεις"

cd Mostlylucid.BlogLLM.Core
dotnet add package LLamaSharp  # Latest version
dotnet add package LLamaSharp.Backend.Cuda12  # Latest, matching CUDA version

Αναζήτηση

Ρύθμιση LLamaSharp

Εγκατάσταση πακέτου NuGetName

using LLama;
using LLama.Common;

// Check if CUDA is available
bool cudaAvailable = NativeLibraryConfig.Instance.CudaEnabled;
Console.WriteLine($"CUDA Available: {cudaAvailable}");

Γιατί δύο πακέτα;false- Κεντρική βιβλιοθήκη

  1. ΚΟΥΔΑConstellation name (optional, probably does not need a translation)
  2. LLamaSharp.Backend.Cuda1212 δυαδικά για επιτάχυνση GPU
  3. Επιβεβαίωση του συστήματος υποστήριξης CUDA

Το LLamaSharp θα ανιχνεύσει αυτόματα το CUDA αν εγκατασταθεί σωστά.

Εάν

using LLama;
using LLama.Common;

namespace Mostlylucid.BlogLLM.Core.Services
{
    public class ModelParameters
    {
        public string ModelPath { get; set; } = string.Empty;
        public int ContextSize { get; set; } = 4096;  // Context window
        public int GpuLayerCount { get; set; } = 35;  // Layers on GPU (35 = all for 7B)
        public int Seed { get; set; } = 1337;  // For reproducibility
        public float Temperature { get; set; } = 0.7f;  // Creativity (0.0 = deterministic, 1.0 = creative)
        public float TopP { get; set; } = 0.9f;  // Nucleus sampling
        public int MaxTokens { get; set; } = 500;  // Max generation length
    }
}

, έλεγχος::

  • **CUDA 12.x εγκατεστημένο (Μέρος 2)**εγκατεστημένο πακέτο

    • Το PATH περιλαμβάνει τον κατάλογο bin CUDA
    • Κατασκευή της υπηρεσίας LLM
  • Μοντέλο παραμέτρωνΕπεξηγήσεις παραμέτρων

    • ΠλαίσιοΜέγεθος
    • : Πόσο κείμενο μπορεί να "δεί" το μοντέλο ταυτόχρονα
    • 4096 μάρκες 3.000 λέξεις
  • Μεγαλύτερο = περισσότερο πλαίσιο αλλά πιο αργό και περισσότερο VRAMGpuLayerCount

    • : Πόσα στρώματα μετασχηματιστή τρέχουν σε GPU
    • 7B μοντέλα έχουν ~32 στρώματα
    • 35 = βάλτε τα πάντα σε GPU (γρήγορο)
  • Χαμηλότερες τιμές = χρήση λιγότερο VRAM αλλά πιο αργήΘερμοκρασία

    • : Έλεγχος της τυχαιότητας
    • 0.0 = πάντα να διαλέγετε το πιο πιθανό σημείο (βαρετό, επαναλαμβανόμενο)

0.7 = καλό υπόλοιπο (η προεπιλογή μας)

using LLama;
using LLama.Common;
using Microsoft.Extensions.Logging;

namespace Mostlylucid.BlogLLM.Core.Services
{
    public interface ILlmService
    {
        Task<string> GenerateAsync(string prompt, CancellationToken cancellationToken = default);
        Task<string> GenerateWithContextAsync(string prompt, List<string> contextChunks, CancellationToken cancellationToken = default);
    }

    public class LlmService : ILlmService, IDisposable
    {
        private readonly LLamaWeights _model;
        private readonly LLamaContext _context;
        private readonly ILogger<LlmService> _logger;
        private readonly ModelParameters _parameters;

        public LlmService(ModelParameters parameters, ILogger<LlmService> logger)
        {
            _parameters = parameters;
            _logger = logger;

            _logger.LogInformation("Loading model from {ModelPath}", parameters.ModelPath);

            // Configure model parameters
            var modelParams = new ModelParams(parameters.ModelPath)
            {
                ContextSize = (uint)parameters.ContextSize,
                GpuLayerCount = parameters.GpuLayerCount,
                Seed = (uint)parameters.Seed,
                UseMemoryLock = true,  // Keep model in RAM
                UseMemorymap = true    // Memory-map the model file
            };

            // Load model
            _model = LLamaWeights.LoadFromFile(modelParams);
            _context = _model.CreateContext(modelParams);

            _logger.LogInformation("Model loaded successfully. VRAM used: ~{VRAM}GB",
                EstimateVRAMUsage(parameters.GpuLayerCount));
        }

        public async Task<string> GenerateAsync(string prompt, CancellationToken cancellationToken = default)
        {
            var executor = new InteractiveExecutor(_context);

            var inferenceParams = new InferenceParams
            {
                Temperature = _parameters.Temperature,
                TopP = _parameters.TopP,
                MaxTokens = _parameters.MaxTokens,
                AntiPrompts = new[] { "\n\nUser:", "###" }  // Stop generation at these
            };

            var result = new StringBuilder();

            _logger.LogInformation("Generating response for prompt: {Prompt}", TruncateForLog(prompt));

            await foreach (var token in executor.InferAsync(prompt, inferenceParams, cancellationToken))
            {
                result.Append(token);
            }

            var response = result.ToString().Trim();
            _logger.LogInformation("Generated {Tokens} tokens", CountTokens(response));

            return response;
        }

        public async Task<string> GenerateWithContextAsync(
            string prompt,
            List<string> contextChunks,
            CancellationToken cancellationToken = default)
        {
            // Build prompt with retrieved context
            var fullPrompt = BuildContextualPrompt(prompt, contextChunks);

            _logger.LogInformation("Context chunks: {Count}, Total prompt tokens: ~{Tokens}",
                contextChunks.Count, CountTokens(fullPrompt));

            return await GenerateAsync(fullPrompt, cancellationToken);
        }

        private string BuildContextualPrompt(string userPrompt, List<string> contextChunks)
        {
            var sb = new StringBuilder();

            sb.AppendLine("You are a helpful writing assistant for a technical blog.");
            sb.AppendLine("Use the following excerpts from past blog posts as context:");
            sb.AppendLine();

            for (int i = 0; i < contextChunks.Count; i++)
            {
                sb.AppendLine($"--- Context {i + 1} ---");
                sb.AppendLine(contextChunks[i]);
                sb.AppendLine();
            }

            sb.AppendLine("---");
            sb.AppendLine();
            sb.AppendLine("Based on the context above, help with the following:");
            sb.AppendLine(userPrompt);
            sb.AppendLine();
            sb.AppendLine("Response:");

            return sb.ToString();
        }

        private int CountTokens(string text)
        {
            // Rough estimate: 1 token ≈ 4 characters
            return text.Length / 4;
        }

        private string TruncateForLog(string text, int maxLength = 100)
        {
            if (text.Length <= maxLength) return text;
            return text.Substring(0, maxLength) + "...";
        }

        private double EstimateVRAMUsage(int gpuLayers)
        {
            // Rough estimate for 7B model
            return (gpuLayers / 35.0) * 6.0;  // ~6GB for full 7B model
        }

        public void Dispose()
        {
            _context?.Dispose();
            _model?.Dispose();
        }
    }
}

1.0+ = πολύ δημιουργικό (μπορεί να είναι αναίσθητο):

  1. TopPunity name (optional, probably does not need a translation): δειγματοληψία Nucleuus
  2. 0.9 = εξετάστε τις μάρκες που αποτελούν το 90% της μάζας πιθανότηταςΑποτρέπει τη δειγματοληψία από πολύ απίθανες μάρκες
  3. Εφαρμογή υπηρεσιών LLMΠώς λειτουργεί
  4. Φόρτωση μοντέλου: Φορτώνει το μοντέλο GGUF σε VRAM χρησιμοποιώντας καθορισμένες παραμέτρους
  5. Διαδραστικός εκτελεστής: Λειτουργία εκτέλεσης LLamaSharp για αλληλεπιδράσεις όπως η συνομιλία

InferAsync

using Microsoft.Extensions.Logging;

class Program
{
    static async Task Main(string[] args)
    {
        // Setup logging
        var loggerFactory = LoggerFactory.Create(builder => builder.AddConsole());
        var logger = loggerFactory.CreateLogger<LlmService>();

        // Configure model
        var parameters = new ModelParameters
        {
            ModelPath = @"C:\models\mistral-7b\mistral-7b-instruct-v0.2.Q5_K_M.gguf",
            ContextSize = 4096,
            GpuLayerCount = 35,
            Temperature = 0.7f,
            MaxTokens = 200
        };

        // Create service
        using var llmService = new LlmService(parameters, logger);

        // Test simple generation
        Console.WriteLine("=== Test 1: Simple Generation ===\n");
        var response1 = await llmService.GenerateAsync(
            "Explain what Docker Compose is in 2-3 sentences."
        );
        Console.WriteLine(response1);
        Console.WriteLine("\n");

        // Test with context
        Console.WriteLine("=== Test 2: Generation with Context ===\n");
        var context = new List<string>
        {
            "Docker Compose is a tool for defining and running multi-container Docker applications. With Compose, you use a YAML file to configure your application's services.",
            "In development, Docker Compose makes it easy to spin up all dependencies (databases, caches, etc.) with one command: docker-compose up."
        };

        var response2 = await llmService.GenerateWithContextAsync(
            "Write an introduction paragraph for a blog post about using Docker Compose for development dependencies.",
            context
        );
        Console.WriteLine(response2);
    }
}

: Streams markes as they're created (real-time expect):

=== Test 1: Simple Generation ===

Docker Compose is a tool that allows you to define and run multi-container Docker applications using a simple YAML configuration file. It simplifies the process of managing multiple containers, networking, and volumes, making it ideal for development environments.

=== Test 2: Generation with Context ===

If you've ever found yourself juggling multiple terminal windows to start databases, caches, and other services for local development, Docker Compose is about to become your new best friend. This powerful tool lets you define your entire development environment in a single YAML file and spin everything up with one command. In this post, we'll explore how to leverage Docker Compose to manage all your development dependencies, making your local setup reproducible, shareable, and incredibly easy to manage.

Κτίριο πλαισίου

: Συνδυάζει τον χρήστη με ανακτημένα κομμάτια blog

Αντι-προβλήματα

: Σταματάει τη γενιά σε ορισμένες χορδές (προφυλάξεις ασυναρτησίες)

namespace Mostlylucid.BlogLLM.Client.Services
{
    public class SuggestionService : ISuggestionService
    {
        private readonly BatchEmbeddingService _embeddingService;
        private readonly QdrantVectorStore _vectorStore;
        private readonly ILlmService _llmService;  // NEW

        public SuggestionService(
            BatchEmbeddingService embeddingService,
            QdrantVectorStore vectorStore,
            ILlmService llmService)  // NEW
        {
            _embeddingService = embeddingService;
            _vectorStore = vectorStore;
            _llmService = llmService;
        }

        public async Task<string> GenerateAiSuggestionAsync(
            string currentText,
            List<SimilarPost> context)
        {
            // Extract text from similar posts
            var contextChunks = context
                .Take(3)  // Top 3 most similar
                .Select(p => p.FullText)
                .ToList();

            // Determine what type of suggestion to generate
            var prompt = DeterminePromptType(currentText);

            // Generate suggestion
            var suggestion = await _llmService.GenerateWithContextAsync(
                prompt,
                contextChunks
            );

            return suggestion;
        }

        private string DeterminePromptType(string currentText)
        {
            // Analyze what user is writing
            var lines = currentText.Split('\n');
            var lastLine = lines.LastOrDefault(l => !string.IsNullOrWhiteSpace(l)) ?? "";

            // Is user starting a new section?
            if (lastLine.StartsWith("## "))
            {
                return "Suggest 3-5 bullet points for what this section could cover.";
            }

            // Is user writing code?
            if (lastLine.Contains("```"))
            {
                return "Suggest relevant code examples that might be useful here.";
            }

            // Is user writing an introduction?
            if (currentText.Length < 500 && currentText.Contains("## Introduction"))
            {
                return "Suggest 2-3 sentences to continue this introduction based on similar posts.";
            }

            // Default: continue current thought
            return "Suggest 1-2 sentences to continue the current paragraph in a natural way.";
        }
    }
}

Έλεγχος της Υπηρεσίας

public partial class SuggestionsViewModel : ViewModelBase
{
    [RelayCommand]
    private async Task RegenerateSuggestion()
    {
        IsGenerating = true;
        AiSuggestion = "Generating...";

        try
        {
            var currentText = GetCurrentEditorText();  // From messaging
            var suggestion = await _suggestionService.GenerateAiSuggestionAsync(
                currentText,
                SimilarPosts.ToList()
            );

            AiSuggestion = suggestion;
        }
        catch (Exception ex)
        {
            AiSuggestion = $"Error: {ex.Message}";
        }
        finally
        {
            IsGenerating = false;
        }
    }
}

Αναμενόμενη έξοδος

Καταπληκτικό!

Το μοντέλο λειτουργεί και παράγει συνεκτικό κείμενο, το οποίο έχει γνώση του πλαισίου.

public class LlmServiceFactory
{
    private static LlmService? _instance;
    private static readonly object _lock = new();

    public static LlmService GetInstance(ModelParameters parameters, ILogger<LlmService> logger)
    {
        if (_instance == null)
        {
            lock (_lock)
            {
                if (_instance == null)
                {
                    _instance = new LlmService(parameters, logger);
                }
            }
        }

        return _instance;
    }
}

Ενσωμάτωση με τον πίνακα προτάσεων

Τώρα ας ενσωματώσουμε την γενιά LLM στον πελάτη των Windows από το Μέρος 5.

public class StatefulLlmService
{
    private readonly InferenceParams _defaultParams;
    private string _cachedPromptPrefix = string.Empty;

    public async Task<string> GenerateWithPrefixAsync(string prefix, string newPrompt)
    {
        // If prefix matches cached, reuse KV cache
        if (prefix == _cachedPromptPrefix)
        {
            // Only process new tokens
            return await GenerateAsync(newPrompt);
        }

        // Process entire prompt and cache
        _cachedPromptPrefix = prefix;
        return await GenerateAsync(prefix + newPrompt);
    }
}

Update ScommodationService

Ενημέρωση προτάσεωνViewModel

Βελτιστοποίηση απόδοσης

public async Task<List<string>> GenerateBatchAsync(List<string> prompts)
{
    var results = new List<string>();

    foreach (var prompt in prompts)
    {
        // With KV cache reuse, subsequent prompts are faster
        results.Add(await GenerateAsync(prompt));
    }

    return results;
}

Μοντέλο Caching

Κρατήστε το μοντέλο φορτωμένο μεταξύ των αιτήσεων:

Επαναχρησιμοποίηση Cache του KVName

private string PromptContinueWriting(string currentText, List<string> context)
{
    return $@"You are a technical blog writing assistant.

Here are excerpts from similar blog posts:
{string.Join("\n\n", context.Select((c, i) => $"--- Post {i + 1} ---\n{c}"))}

Current draft:
{currentText}

Task: Suggest 2-3 sentences to naturally continue the current paragraph.
Keep the same technical depth and casual, pragmatic tone.

Suggestion:";
}

Η LLamaSharp υποστηρίζει την επαναχρησιμοποίηση του KV cache για γρηγορότερες επόμενες γενιές:

private string PromptSectionStructure(string sectionTitle, List<string> context)
{
    return $@"You are a technical blog writing assistant.

Similar sections from past posts:
{string.Join("\n\n", context)}

New section: {sectionTitle}

Task: Suggest 4-6 bullet points for what this section should cover.
Format as a markdown list.

Bullets:";
}

Αυτό είναι ιδιαίτερα χρήσιμο για τη χρήση μας περίπτωση - το πλαίσιο κομμάτια παραμένουν το ίδιο, μόνο η ερώτηση του χρήστη αλλάζει.

private string PromptCodeExample(string description, List<string> context)
{
    return $@"You are a C# coding assistant.

Relevant code from past posts:
{string.Join("\n\n", context)}

Task: {description}

Provide a clean, well-commented C# code example.

Code:";
}

Παρτίδα

Για πολλαπλές προτάσεις, τα παρτίδα:

Prompt Engineering for Writing Assistance

public class VramMonitor
{
    [DllImport("nvml.dll")]
    private static extern int nvmlDeviceGetMemoryInfo(IntPtr device, ref NvmlMemory memory);

    [StructLayout(LayoutKind.Sequential)]
    public struct NvmlMemory
    {
        public ulong Total;
        public ulong Free;
        public ulong Used;
    }

    public static (ulong used, ulong total) GetVramUsage()
    {
        // Simplified - actual implementation needs proper NVML initialization
        var memory = new NvmlMemory();
        // nvmlDeviceGetMemoryInfo(device, ref memory);

        return (memory.Used / 1024 / 1024, memory.Total / 1024 / 1024);  // Convert to MB
    }
}

Καλή ώθηση = καλή έξοδος.

public class LlmServiceWithUnload : IDisposable
{
    private LlmService? _service;
    private readonly Timer _unloadTimer;
    private DateTime _lastUsed;

    public LlmServiceWithUnload()
    {
        _unloadTimer = new Timer(CheckForUnload, null, TimeSpan.FromMinutes(1), TimeSpan.FromMinutes(1));
    }

    private void CheckForUnload(object? state)
    {
        if (_service != null && (DateTime.Now - _lastUsed) > TimeSpan.FromMinutes(10))
        {
            _service.Dispose();
            _service = null;
            GC.Collect();
            Console.WriteLine("Model unloaded due to inactivity");
        }
    }

    public async Task<string> GenerateAsync(string prompt)
    {
        _lastUsed = DateTime.Now;

        if (_service == null)
        {
            // Reload model
            _service = CreateService();
        }

        return await _service.GenerateAsync(prompt);
    }
}

Εδώ είναι τα πρότυπα για διάφορα σενάρια:

Συνεχίστε να γράφετε

public async Task<string> GenerateWithRetryAsync(string prompt, int maxRetries = 3)
{
    for (int i = 0; i < maxRetries; i++)
    {
        try
        {
            return await GenerateAsync(prompt);
        }
        catch (OutOfMemoryException)
        {
            _logger.LogWarning("OOM error, reducing max tokens");
            _parameters.MaxTokens = Math.Max(100, _parameters.MaxTokens / 2);
        }
        catch (Exception ex)
        {
            _logger.LogError(ex, "Generation failed, attempt {Attempt}/{Max}", i + 1, maxRetries);

            if (i == maxRetries - 1) throw;

            await Task.Delay(1000 * (i + 1));  // Exponential backoff
        }
    }

    throw new Exception("Generation failed after retries");
}

Προτείνετε τη δομή τμήματος

Παράδειγμα κώδικα

  1. ✅ Chose Διαχείριση μνήμηςΜε μεγάλα μοντέλα, η διαχείριση μνήμης είναι ζωτικής σημασίας.
  2. ✅ Understood Παρακολούθηση χρήσης VRAMΑποφόρτωση μοντέλου όταν δεν είναι ενεργό
  3. ✅ Selected appropriate model (Χειρισμός λάθους / Τα LLM μπορούν να αποτύχουν με απροσδόκητους τρόπους.Χειριστείτε με χάρη:
  4. ✅ Implemented LlmService with CUDA acceleration
  5. ✅ Integrated with Windows client for suggestions
  6. ✅ Implemented prompt engineering for writing tasks
  7. ✅ Added performance optimizations (caching, batching)
  8. ✅ Handled memory management and errors

Περίληψη

Έχουμε ενσωματώσει με επιτυχία το τοπικό συμπέρασμα LLM:**LLamaSharpCity name (optional, probably does not need a translation)**για την ενσωμάτωση C#

  • Μορφή GGUF
  • και ποσοτικοποίηση
  • Mistral 7B
  • Λάμα 3
  • 8B)
  • Ποιο είναι το επόμενο;
  • Το

Μέρος 7: Παραγωγή Περιεχομένου & Prompt Engineering

, θα επικεντρωθούμε στον πλήρη αγωγό παραγωγής περιεχομένου:

Μέρος 1: Εισαγωγή & Αρχιτεκτονική

Μέρος 6: Τοπική ενσωμάτωση LLM(Παρούσα θέση)!

logo

© 2026 Scott Galloway — Unlicense — All content and source code on this site is free to use, copy, modify, and sell.