1 matches found
Secure Speculative Decoding for Large Language Models
Speculative decoding accelerates inference for a large language model LLM, referred to as the target model, by first using a smaller model, referred to as the draft model, to generate candidate tokens and then verifying them with the target model for acceptance or rejection. Prior studies primari...
SaveExploits0
20