3 matches found
Agentic-CLIP-Benchmark
Agentic-CLIP-Benchmark: 在CIFAR-10上的零样本评估 -green.svg 本项目实现了一个稳健的自动化管道,用于在完整的 CIFAR-10 测试集(10,000张图片)上评估OpenAI的 CLIP ViT-B/32 模型。采用 AI原生工作流(Trae IDE) 开发,实现了高精度的零样本准确率 88.80% 。 📊 性能总结 Top-1准确率 : 88.80% 零样本 模型 : openai/clip-vit-base-patch32 使用Safetensors 推理硬件 : NVIDIA RTX 2060 时间复杂度 : 通过批推理优化(批大小:32)...
Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts
The deployment of large language models LLMs in Swiss financial and regulatory contexts demands empirical evidence of both production reliability and adversarial security, dimensions not jointly operationalized in existing Swiss-focused evaluation frameworks. This paper introduces Swiss-Bench 003...
SecureRAG-RTL: A Retrieval-Augmented, Multi-Agent, Zero-Shot LLM-Driven Framework for Hardware Vulnerability Detection
Large language models LLMs have shown remarkable capabilities in natural language processing tasks, yet their application in hardware security verification remains limited due to scarcity of publicly available hardware description language HDL datasets. This knowledge gap constrains LLM performan...