This is a viewer only at the moment see the article on how this works.
To update the preview hit Ctrl-Alt-R (or ⌘-Alt-R on Mac) or Enter to refresh. The Save icon lets you save the markdown file to disk
This is a preview from the server running through my markdig pipeline
Monday, 24 November 2025
Vuoi ottenere bel, descrittivo alt testo per le immagini sui vostri siti o jsut estrarre testo da loro? mostlylucid.llmalttext utilizza il modello di linguaggio di visione Florence-2 di Microsoft per generare testo alt di alta qualità automaticamente - in esecuzione interamente localmente sulla vostra macchina, nessuna chiave API richiesta.
Nota: Ho bisogno di aggiornare questo documento ora il pacchetto nuget è fuori. Se si guarda qui Lei finirà un sito di dimostrazione nifty che Lei può scaricare e usare. Lo aggiornerò con i dettagli nei prossimi giorni.
Alt testo conta. I lettori di schermo dipendono da esso, fattori di classifica SEO, ed è semplicemente la cosa giusta da fare per l'accessibilità. Ma la scrittura di un buon testo alt per centinaia di immagini? Ecco dove la maggior parte di noi non riescono.
Questo pacchetto risolve il problema utilizzando il modello di linguaggio di visione Florence-2 di Microsoft - eseguito interamente localmente sulla macchina, nessuna chiave API richiesta.
Codice sorgente: github.com/scottgal/mostlylucid.nugetpackages
Ogni <img> tag dovrebbe avere un testo alt significativo. Ma in pratica:
Che cosa succede se si può generare il testo di alta qualità automaticamente, in esecuzione interamente sul proprio hardware?
Il pacchetto utilizza il modello Florence-2 di Microsoft tramite runtime ONNX. Ecco la pipeline di elaborazione:
flowchart TB
subgraph Input[Image Sources]
A[File Path]
B[URL]
C[Stream]
D[Byte Array]
end
subgraph Processing[Florence-2 Pipeline]
E[Image Preprocessing]
F[Vision Encoder]
G[Language Decoder]
end
subgraph Output[Results]
H[Alt Text]
I[OCR Text]
J[Content Type]
end
A --> E
B --> E
C --> E
D --> E
E --> F
F --> G
G --> H
G --> I
G --> J
style A stroke:#10b981,stroke-width:2px
style B stroke:#10b981,stroke-width:2px
style C stroke:#10b981,stroke-width:2px
style D stroke:#10b981,stroke-width:2px
style F stroke:#6366f1,stroke-width:2px
style G stroke:#6366f1,stroke-width:2px
style H stroke:#ec4899,stroke-width:2px
style I stroke:#ec4899,stroke-width:2px
style J stroke:#ec4899,stroke-width:2px
Caratteristiche principali:
dotnet add package Mostlylucid.LlmAltText
// Program.cs
builder.Services.AddAltTextGeneration();
La prima esecuzione scarica il modello Florence-2 (~800MB), poi sei pronto a partire.
public class ImageController : ControllerBase
{
private readonly IImageAnalysisService _imageAnalysis;
public ImageController(IImageAnalysisService imageAnalysis)
{
_imageAnalysis = imageAnalysis;
}
[HttpPost("analyze")]
public async Task<IActionResult> Analyze(IFormFile image)
{
using var stream = image.OpenReadStream();
var altText = await _imageAnalysis.GenerateAltTextAsync(stream);
return Ok(new { altText });
}
}
Il servizio accetta immagini da qualsiasi luogo - file, URL, flussi o array di byte.
var altText = await _imageAnalysis.GenerateAltTextFromFileAsync("/images/photo.jpg");
var altText = await _imageAnalysis.GenerateAltTextFromUrlAsync(
"https://example.com/image.png");
using var stream = file.OpenReadStream();
var altText = await _imageAnalysis.GenerateAltTextAsync(stream);
var bytes = await httpClient.GetByteArrayAsync(imageUrl);
var altText = await _imageAnalysis.GenerateAltTextAsync(bytes);
Florence-2 supporta tre modalità di didascalia. Scegli in base alle tue esigenze:
// Brief - "A dog sitting on grass"
var brief = await _imageAnalysis.GenerateAltTextAsync(stream, "CAPTION");
// Detailed - "A golden retriever sitting on green grass in a park"
stream.Position = 0;
var detailed = await _imageAnalysis.GenerateAltTextAsync(stream, "DETAILED_CAPTION");
// Most detailed (default) - Full accessibility description
stream.Position = 0;
var full = await _imageAnalysis.GenerateAltTextAsync(stream, "MORE_DETAILED_CAPTION");
// "A happy golden retriever with light fur sitting on lush green grass
// in a sunny park, with trees visible in the background."
Quando usare ciascuna:
| Tipo di operazione | Migliore per |
|---|---|
CAPTION |
Miniature, immagini decorative, suggerimenti rapidi |
DETAILED_CAPTION |
Social media, accessibilità di base |
MORE_DETAILED_CAPTION |
Piena accessibilità, lettori dello schermo (raccomandati) |
Firenze-2 può anche estrarre testo da immagini - utili per screenshot, documenti e grafici.
// Extract text only
var extractedText = await _imageAnalysis.ExtractTextAsync(stream);
// Get both alt text and extracted text
var (altText, ocrText) = await _imageAnalysis.AnalyzeImageAsync(stream);
Console.WriteLine($"Alt: {altText}");
Console.WriteLine($"OCR: {ocrText}");
Non tutte le immagini sono uguali. Una fotografia ha bisogno di un testo descrittivo; un documento ha bisogno del suo contenuto di testo. La funzione di classificazione ti aiuta a gestire ciascuno in modo appropriato:
var result = await _imageAnalysis.AnalyzeWithClassificationAsync(stream);
Console.WriteLine($"Type: {result.ContentType}"); // e.g., "Photograph"
Console.WriteLine($"Confidence: {result.ContentTypeConfidence:P0}"); // e.g., "87%"
Console.WriteLine($"Has Text: {result.HasSignificantText}");
var result = await _imageAnalysis.AnalyzeWithClassificationAsync(stream);
switch (result.ContentType)
{
case ImageContentType.Document:
// Documents - prioritize extracted text
return result.ExtractedText;
case ImageContentType.Screenshot:
// Screenshots - combine description with UI text
return result.HasSignificantText
? $"{result.AltText}. Text visible: {result.ExtractedText}"
: result.AltText;
case ImageContentType.Chart:
// Charts - describe the visualization plus data
return $"{result.AltText}. Data: {result.ExtractedText}";
case ImageContentType.Photograph:
default:
// Photos - just the description
return result.AltText;
}
| Tipo | Descrizione | Esempio |
|---|---|---|
Photograph |
Foto del mondo reale | Persone, paesaggi, prodotti |
Document |
Contenuti pesanti per testo | PDF, moduli, articoli |
Screenshot |
Cattura software | UI, siti web, applicazioni |
Chart |
Visualizzazioni dati | Grafici, grafici a torta, tabelle |
Illustration |
Contenuto di disegno | Opere, cartoni animati, icone |
Diagram |
Disegni tecnici | Carrelli di flusso, UML, schemi |
Unknown |
Non classificati | Casi di bordi |
Qui è dove diventa interessante. Il TagHelper genera automaticamente il testo dell'alt per qualsiasi <img> tag mancante uno - al momento del rendering.
// Program.cs
builder.Services.AddAltTextGeneration(options =>
{
options.EnableTagHelper = true;
options.EnableDatabase = true; // Cache results
options.DbProvider = AltTextDbProvider.Sqlite;
options.SqliteDbPath = "./alttext.db";
});
var app = builder.Build();
await app.Services.MigrateAltTextDatabaseAsync();
Registrare il tagHelper in _ViewImports.cshtml:
@addTagHelper *, Mostlylucid.LlmAltText
flowchart LR
subgraph Razor[Razor View Rendering]
A[img tag found]
B{Has alt attribute?}
C[Skip - use existing]
D{In cache?}
E[Return cached]
F[Fetch image]
G[Generate alt text]
H[Cache result]
I[Render with alt]
end
A --> B
B -->|Yes| C
B -->|No| D
D -->|Yes| E
D -->|No| F
F --> G
G --> H
H --> I
E --> I
style A stroke:#10b981,stroke-width:2px
style B stroke:#6366f1,stroke-width:2px
style G stroke:#ec4899,stroke-width:2px
style I stroke:#8b5cf6,stroke-width:2px
<!-- NO ALT - Will be processed -->
<img src="https://example.com/photo.jpg" />
<!-- HAS ALT - Skipped (respects your text) -->
<img src="https://example.com/photo.jpg" alt="My custom description" />
<!-- EMPTY ALT - Skipped (decorative image per a11y standards) -->
<img src="https://example.com/decorative.jpg" alt="" />
<!-- EXPLICIT SKIP - Skipped -->
<img src="https://example.com/photo.jpg" data-skip-alt="true" />
<!-- DATA URI - Skipped (can't fetch) -->
<img src="data:image/png;base64,..." />
<!-- RELATIVE PATH - Skipped (needs absolute URL) -->
<img src="/images/photo.jpg" />
Per la sicurezza, è possibile limitare i domini che il TagHelper prenderà da:
options.AllowedImageDomains = new List<string>
{
"mycdn.example.com",
"images.mysite.org",
"cdn.githubusercontent.com"
};
Senza cache, ogni rendering di pagina rigenererebbe il testo alt. Questo è lento e sprecoso. Il deposito della cache del database risulta chiavi in mano dall'URL dell'immagine.
builder.Services.AddAltTextGeneration(options =>
{
options.EnableDatabase = true;
options.DbProvider = AltTextDbProvider.Sqlite;
options.SqliteDbPath = "./alttext.db";
options.CacheDurationMinutes = 60;
});
builder.Services.AddAltTextGeneration(options =>
{
options.EnableDatabase = true;
options.DbProvider = AltTextDbProvider.PostgreSql;
options.ConnectionString = Configuration.GetConnectionString("AltTextDb");
});
builder.Services.AddAltTextGeneration(options =>
{
// Model location (~800MB downloaded here)
options.ModelPath = "./models";
// Default task type for alt text generation
options.DefaultTaskType = "MORE_DETAILED_CAPTION";
// Maximum word count for alt text
options.MaxWords = 90;
// Enable detailed logging
options.EnableDiagnosticLogging = true;
// TagHelper settings
options.EnableTagHelper = true;
options.EnableDatabase = true;
options.AutoMigrateDatabase = true;
// Database provider
options.DbProvider = AltTextDbProvider.Sqlite;
options.SqliteDbPath = "alttext.db";
// or
options.DbProvider = AltTextDbProvider.PostgreSql;
options.ConnectionString = "Host=localhost;Database=alttext;...";
// Security
options.AllowedImageDomains = new List<string> { "cdn.example.com" };
options.SkipSrcPrefixes = new List<string> { "data:", "blob:" };
// Caching
options.CacheDurationMinutes = 60;
});
Ecco come lo uso per elaborare le immagini durante l'importazione dei post del blog:
public class ImageProcessor
{
private readonly IImageAnalysisService _imageAnalysis;
private readonly ILogger<ImageProcessor> _logger;
public ImageProcessor(
IImageAnalysisService imageAnalysis,
ILogger<ImageProcessor> logger)
{
_imageAnalysis = imageAnalysis;
_logger = logger;
}
public async Task ProcessMarkdownImagesAsync(string markdownPath)
{
var imageDir = Path.Combine(Path.GetDirectoryName(markdownPath)!, "images");
if (!Directory.Exists(imageDir)) return;
var images = Directory.GetFiles(imageDir, "*.*")
.Where(f => IsImageFile(f));
foreach (var imagePath in images)
{
try
{
var result = await _imageAnalysis
.AnalyzeWithClassificationFromFileAsync(imagePath);
_logger.LogInformation(
"Processed {File}: {Type} ({Confidence:P0})",
Path.GetFileName(imagePath),
result.ContentType,
result.ContentTypeConfidence);
// Store alt text for later use
await SaveAltTextAsync(imagePath, result.AltText);
}
catch (Exception ex)
{
_logger.LogWarning(ex, "Failed to process {File}", imagePath);
}
}
}
private static bool IsImageFile(string path)
{
var ext = Path.GetExtension(path).ToLowerInvariant();
return ext is ".jpg" or ".jpeg" or ".png" or ".gif" or ".webp" or ".bmp";
}
}
| Metric | Tipic Value |
|---|---|
| Prima esecuzione | Più lento (download del modello ~800MB) |
| Carico del modello | 1-3 secondi |
| Lavorazione per immagine | 500-2000ms |
| Uso della memoria | 2GB+ raccomandato |
| Spazio su disco | ~ 800MB per i modelli |
// 1. Register as Singleton (model load is expensive)
builder.Services.AddAltTextGeneration(); // Already singleton internally
// 2. Check readiness before processing
if (!_imageAnalysis.IsReady)
{
return StatusCode(503, "AI model still initializing");
}
// 3. Use cancellation tokens for timeouts
var cts = new CancellationTokenSource(TimeSpan.FromSeconds(30));
var altText = await _imageAnalysis.GenerateAltTextFromUrlAsync(url, cts.Token);
// 4. Process in batches, not parallel (memory constraints)
foreach (var image in images)
{
await ProcessImageAsync(image); // Sequential is safer
}
Il pacchetto include il tracciamento integrato:
builder.Services.AddOpenTelemetry()
.WithTracing(tracing =>
{
tracing.AddSource("Mostlylucid.LlmAltText");
});
Attività rintracciate:
llmalttext.generate_alt_textllmalttext.extract_textllmalttext.analyze_imagellmalttext.classify_content_typeAggiungi un controllo dello stato di salute per monitorare lo stato del modello:
public class AltTextHealthCheck : IHealthCheck
{
private readonly IImageAnalysisService _service;
public AltTextHealthCheck(IImageAnalysisService service)
=> _service = service;
public Task<HealthCheckResult> CheckHealthAsync(
HealthCheckContext context,
CancellationToken cancellationToken = default)
{
return Task.FromResult(_service.IsReady
? HealthCheckResult.Healthy("Florence-2 model ready")
: HealthCheckResult.Unhealthy("Model not initialized"));
}
}
// Registration
builder.Services.AddHealthChecks()
.AddCheck<AltTextHealthCheck>("alttext");
Error: Failed to download model files
Soluzioni:
ModelPath_imageAnalysis.IsReady // Returns false
Soluzioni:
Soluzioni:
MORE_DETAILED_CAPTION (predefinito)Soluzioni:
EnableTagHelper = true@addTagHelper in _ViewImports.cshtmlAllowedImageDomains configurazioneIl testo alt generato è un punto di partenza. Per i migliori risultati:
alt="" per immagini puramente decorativeMostlylucid.LlmAltText porta l'accessibilità AI-powered alle vostre applicazioni .NET senza il costo o problemi di privacy di API esterne. Il TagHelper rende particolarmente facile - basta abilitare e il vostro <img> tags ottiene testo automatico alt.
Il pacchetto è Unlicense (public domain), quindi fai quello che vuoi.
© 2026 Scott Galloway — Unlicense — All content and source code on this site is free to use, copy, modify, and sell.