llama.cpp will not refuse a k-quant or i-quant whose first tensor dimension is not divisible by 256. tensor_type_fallback swaps that tensor to a block-32 type and logs a warning. IQ2_XXS, IQ2_XS, IQ2_S, IQ3_XXS, IQ3_S, and IQ4_XS tensors become IQ4_NL. Q2_K and Q3_K tensors become Q4_0. Q6_K tensors become Q8_0. If the replacement still does not divide, it falls through to F16. Other tensors in the same file can keep the requested type. The file keeps the name you asked for.
NVIDIA’s Nemotron-3.5-Lightning config has hidden_size 2688 and moe_intermediate_size 1856. 2688 % 256 = 128. 1856 % 256 = 64.
On bartowski’s GGUFs, the four IQ2 rungs are 18.84–18.85 GB. IQ4_NL is 18.92 GB. Q6_K is 34.31 GB next to Q8_0 at 35.00 GB. Those sizes are a mix receipt, not a proof that every tensor landed as the fallback.
I would score {requested, landed, bpw} per tensor. A filename is the requested quant, not the landed type. I did not dump tensor types from the files.