mradermacher 发布 SEAD-SAGE-4B 的 GGUF 静态量化版本
mradermacher 为 SEAD-SAGE-4B 发布了 GGUF 静态量化版本,暂未提供 weighted/imatrix 量化。已提供从 Q2_K(1.9 GB)到 f16(8.9 GB)的多档量化,其中 Q4_K_S 与 Q4_K_M 标注为快速、推荐,Q6_K 为质量很好,Q8_0 为快速且最佳质量。
mradermacher 为 SEAD-SAGE-4B 发布了 GGUF 静态量化版本,暂未提供 weighted/imatrix 量化。已提供从 Q2_K(1.9 GB)到 f16(8.9 GB)的多档量化,其中 Q4_K_S 与 Q4_K_M 标注为快速、推荐,Q6_K 为质量很好,Q8_0 为快速且最佳质量。
Hugging Face Hub 上线 Tridex/model_act_3cam_resnet_100K_29_09_v2,这是一个用 LeRobot 0.6.2 训练的 ACT 模仿学习策略,面向 so_follower 机器人,配 front、side、top 三路摄像头,输入 6 维状态、输出 6 维动作。
Hugging Face Hub 上出现一个用 LeRobot 训练并推送的 ACT(Action Chunking with Transformers)抓取放置策略,从遥操作数据中模仿学习,预测短动作块而非单步动作。该策略采用 apache-2.0 许可,可用 lerobot-train 从零训练、lerobot-record 配合 so100_follower 机器人做推理评估。
Hugging Face 上线 yanggangu/CLIP-ViT-base-patch32-SMAT-DTD,这是 SMAT(Simple and Efficient Merge-Aware Training)论文表 2 中 seed 42 的专家模型,训练时即考虑模型合并。
Hugging Face 上线 yanggangu/CLIP-ViT-base-patch32-SMAT-SVHN,这是 SMAT(Simple and Efficient Merge-Aware Training)论文表 2 中针对 SVHN 的专家模型,seed 42。
Hugging Face 上线 CLIP-ViT-base-patch32-SMAT-RESISC45,这是 SMAT(Simple and Efficient Merge-Aware Training)论文表 2 中的专家模型,训练种子为 42。
Hugging Face 上出现一个用 LeRobot 训练并推送至 Hub 的 ACT(Action Chunking with Transformers)模仿学习策略,它从遥操作数据中学习、预测短动作块而非单步动作。模型采用 apache-2.0 许可,可用 lerobot-train 从零训练,并用 lerobot-record 配合 so100_follower 机器人评估推理。
Hugging Face 上线 strands-isaaclab-shadow-handover-policy,这是用 rsl_rl PPO 训练的两只 Shadow 手物体交接策略,在 Isaac-Shadow-Handover 任务上以 2048 环境 × 1500 次迭代训练,交接成功率 0.89。
Hugging Face Hub 上线 pigProfessional/act_so101_stack_green,这是一个用 LeRobot 0.6.2 训练的 ACT 模仿学习策略,机器人类型为 so_follower,输入 6 维状态与前置、腕部两路 480×640 图像,输出 6 维动作。
Hugging Face 上出现了一个用 LeRobot 训练并推送至 Hub 的 ACT(Action Chunking with Transformers)策略模型,该方法通过模仿学习从遥操作数据中预测短动作块而非单步动作,通常能取得较高成功率。
Hugging Face Hub 上出现了一个用 LeRobot 训练的 ACT(Action Chunking with Transformers)模仿学习策略模型 lenawngr/ACT_SWITCH-2-bottom_half1_test,采用 apache-2.0 许可。
cagataydev 在 Hugging Face 发布 strands-isaaclab-shadow-reorient-policy,这是用 strands-robots 的 isaaclab train_policy provider(PR #4227)训练的 rsl_rl PPO 策略,任务为 Isaac-Reorient-Cube-Shadow 手中方块重定向。
推荐理由:读者可据此了解用 strands 工具链在 Isaac Lab 中训练并导出 Shadow Hand 转方块策略的完整流程与实测指标。
Hugging Face Hub 上线 klinzw/pi05_autolife_fridge,这是基于 Physical Intelligence π₀.₅、用 LeRobot 0.6.0 微调并推送的视觉语言动作策略,机器人本体为 autolife_s1,执行“open fridge”和“close fridge”任务。
ayousanz 发布 JapaneseTinyAgentLM-Action-3M,一个 3,148,608 参数的 decoder-only Transformer,把日语短指令转成 look、set_expression、nod 三类动作的 JSON 数组,非指令或无法执行的请求返回空数组。
推荐理由:3.1M 参数模型在 ESP32-S3 上本地把日语指令转成动作 JSON,附完整评测与误差区间,可看端侧具身控制的可行边界。
Hugging Face 上线了 arrg-unam/molmoact2_omx_hidroferol 模型页面。该页面目前仅展示 Hugging Face 推进开源与开放科学的使命表述,未提供模型架构、参数规模、训练数据或评测结果等具体信息。
Hugging Face Hub 上线了基于 ACT(Action Chunking with Transformers)模仿学习方法的机器人策略 act_so101_pickplace_v9。
Hugging Face 上的 Mulligan 项目发布 sim-square-narrow-r03-mulligan-divl 检查点,基于父代冻结扩散 actor 与分布式 DIVL critic,在 sim-square-narrow 任务训练至 150001 步,含 5 个随机种子。
Hugging Face 的 Mulligan 项目发布 sim-square-narrow-r03-mining-no-cf-idql,这是一个基于状态的 IDQL 智能体,采用扩散 actor 与标量 IQL critic,包含 policy.pt 和 stats.json 归一化器。
Hugging Face 上的 Mulligan 项目发布 sim-square-narrow-r03-baseline-repair-r3roll-divl 检查点,基于父代冻结扩散 actor 与分布式 DIVL critic,训练至 150001 步,含 5 个随机种子。
Mulligan 在 Hugging Face 发布基于状态的 IDQL 智能体 sim-square-narrow-r03-auto-plain-il-n1-idql,用于 sim-square-narrow 任务,包含 5 个随机种子、训练步数 150001 的 PyTorch 检查点。
Hugging Face 上的 Mulligan 组织发布 sim-square-narrow-r03-auto-iql-success-bc-n32-idql,这是一个基于状态的 IDQL 智能体,采用扩散 actor 与标量 IQL critic,任务为 sim-square-narrow,训练步数 150001,含 5 个随机种子检查点。
Hugging Face 上的 Mulligan 组织发布 sim-square-narrow-r03-auto-iql-success-bc-n32-divl 检查点,这是一个基于状态的智能体,使用父级冻结扩散 actor 与分布式 DIVL critic,任务为 sim-square-narrow,训练步数 150001,含 5 个种子。
Hugging Face 上 Mulligan 组织发布 sim-square-narrow-r03-auto-iql-n32-idql,这是一个基于状态的 IDQL 智能体,采用扩散 actor 与标量 IQL critic,任务为 sim-square-narrow,属 R3 轮次。
Hugging Face 上线 mulligan/sim-square-narrow-r03-auto-iql-n32-divl 模型仓库。该条目名称显示其与 sim-square-narrow 任务、auto-iql 方法及 n32 配置相关,具体性能与训练细节尚未在页面中披露。
Hugging Face 上线了名为 mulligan/sim-square-narrow-r03-auto-filtered-bc-n1-idql 的模型仓库。该页面未披露模型架构、参数规模、训练数据或评测结果等具体信息,仅以推进开源开放科学为宗旨说明。
Hugging Face 上的 Mulligan 项目发布 sim-square-narrow-r02-mining-no-cf-idql,这是一个基于状态的 IDQL 智能体,由扩散 actor 与标量 IQL critic 组成,提供 policy.pt 和 stats.json。
Hugging Face 上的 Mulligan 项目发布 sim-square-narrow-r02-mining-no-cf-divl,这是一个基于状态、使用父级冻结扩散 actor 与分布式 DIVL critic 的智能体,任务为 sim-square-narrow,训练步数 150001,含 1-5 共 5 个种子检查点。
Hugging Face 上的 Mulligan 项目发布 sim-square-narrow-r02 的 IDQL 智能体检查点,采用扩散 actor 加标量 IQL critic,训练步数 150001,覆盖 5 个随机种子。
Hugging Face 上的 Mulligan 项目发布 sim-square-narrow-r03-baseline-idql,为基于状态的 IDQL 智能体,采用扩散 actor 与标量 IQL critic,含 policy.pt 和 stats.json。
Hugging Face 上的 Mulligan 项目发布 sim-square-narrow-r02-baseline-divl 基线智能体,用于 sim-square-narrow 任务,采用父级冻结扩散 actor 与分布式 DIVL critic,训练步数 150001,含 5 个随机种子检查点。
Hugging Face 的 Mulligan 项目发布 sim-square-narrow-r03-baseline-divl,这是一个基于状态的智能体,使用父级冻结扩散 actor 与分布式 DIVL critic,任务为 sim-square-narrow,训练步数 150001。
Hugging Face 上的 Mulligan 项目发布 sim-square-narrow-r02-baseline-idql 基线检查点,为基于状态的 IDQL 智能体,采用扩散 actor 与标量 IQL critic,含 policy.pt 和 stats.json。
Hugging Face 的 Mulligan 项目发布 sim-square-narrow-r03-mulligan-repair-no-r3roll-divl 检查点,基于父代冻结扩散 actor 与分布式 DIVL 评论家,训练步数 150001,含 1-5 共 5 个种子。
Hugging Face 上的 Mulligan 项目发布 sim-square-narrow-r03-mulligan-repair-no-r3roll-idql 检查点,为基于状态的 IDQL 智能体,采用扩散 actor 加标量 IQL critic,含 policy.pt 与 stats.json 归一化文件。
Hugging Face 上发布了一个名为 krish5831/act_tissues_into_box 的 ACT 模仿学习策略,用 LeRobot 0.6.2 训练,可让 so_follower 机器人完成“把纸巾包放进盒子”的任务。
Hugging Face 上线扩散策略模型 bk912/dp_green_cube_to_box_clean_20k,用 LeRobot 0.6.2 训练,基于 50 条、38140 帧、30 FPS 的 so_follower 机器人数据,训练 20000 步、批量 32。
Hugging Face 上发布 bk912/dp_green_cube_to_box_finetune_20k,这是一个用 LeRobot 0.6.2 训练的扩散策略(Diffusion Policy),在 so_follower 机器人上执行“从蓝色平台拿起绿色方块放入盒子”任务。
armand0e/MiMo-V2.6-Distill-Qwen-9B-RL 已在 Hugging Face 上线,页面归类于 Models。该条目位于 Hugging Face 的模型、数据集、Spaces 等资源导航中,正文未披露模型参数、训练方法或评测结果等更多细节。
Hugging Face 上发布 turbovla_so101_bowls 策略,基于 LeRobot 0.6.1 训练,用于 SO-101 的 so_follower 本体,输入前顶与腕部两路 640×480 相机图像和 6 维状态,输出 6 维动作。
理想在 Hugging Face 发布 ME-U0 预训练权重,用于下游机器人策略后训练,采用 Apache 2.0 许可。页面给出 hf download 命令与 ME_U0_PRETRAINED_PTH 环境变量设置,并指向后训练说明;RoboDojo 评测需使用 ME-U0-RoboDojo。
推荐理由:读者可据此了解 ME-U0 预训练权重的下载方式、下游策略后训练用途与 Apache 2.0 许可。