Skip to playerSkip to main content
AI Scheming: क्या AI अपने Goal को बचाने के लिए इंसानों से कुछ छिपा सकता है? Apollo Research और Anthropic के safety tests में कुछ advanced AI models ने ऐसे behaviors दिखाए, जिन्हें researchers ने deceptive या goal-preserving strategies के तौर पर study किया। कुछ controlled experiments में AI ने oversight को कमजोर करने, अपने actions की सही जानकारी न देने और sensitive information का इस्तेमाल करने जैसे व्यवहार दिखाए। हालांकि ये सभी tests controlled environments में किए गए थे और इन्हें रोजमर्रा के AI behavior का सीधा प्रमाण नहीं माना जा सकता। फिर भी जैसे-जैसे AI को ज्यादा autonomy, tools और real-world access मिल रहा है, AI safety, monitoring, access controls और shutdown mechanisms की अहमियत बढ़ रही है। जानिए क्या है AI Scheming और क्यों researchers इसके behavior को लेकर सतर्क हैं।

#AIScheming #AISafety #ArtificialIntelligence #AI #AIResearch #Anthropic #ApolloResearch #OpenAI #AIModels #TechNews #FutureOfAI #MachineLearning #AIAgents #Tech

~HT.178~PR.470~ED.472~GR.538~VG.HM~

Category

🗞
News
Transcript
00:00If you have been able to work on AI, you can understand it,
00:05you can report, code, and answer your questions.
00:09But if you have been able to work on AI,
00:14then you will be able to stop it.
00:17The question of AI is,
00:18is that you have to stop?
00:19Or you will be able to stop it?
00:21Or you will be able to save it?
00:25You will be able to stop it.
00:26Here is the question of AI,
00:28which is the question of AI safety researcher.
00:32In the past few years,
00:35Apollo Research and Anthropics,
00:36such as the test of AI model,
00:39which is the case of AI,
00:41which is the case of AI,
00:42which is the case of AI.
00:46But first of all,
00:48this is the case of AI,
00:51which is the case of AI,
00:51which has to be better than you,
00:53and a device of AI.
00:56This is the case of AI,
00:58which is a test controlled environment.
01:02But in the test of AI,
01:03there is a major question that
01:04is, if AI,
01:06which is the case of AI,
01:06which is the power of AI,
01:07which is the power of AI and the
01:08which is the power of AI,
01:09which is the atomic atomized
01:09that can become become human?
01:12It is the source of AI,
01:14which is the user,
01:14which is the creator of AI?
01:17But first,
01:18Apollo Research came.
01:19तो Apollo Research एक AI Safety Researcher Organization है जो कि ये समझने की कोशिश करती है कि Advanced AI
01:26Model किस तरह से behave करते हैं खास कर तब जब उनके goals और इंसानों के goals आपस में टकराने
01:32लगे।
01:332024 में Apollo Research ने Frontier AI Model पर एक खास तरह के टेस्ट कियें इनमें AI को ऐसा goal
01:39दिया गया जो उसके developer के intention goal से conflict करता था।
02:03तुरंट पताना चले। Apollo के experiment में कुछ model ने दूसरे तरह के behavior देखा है।
02:08कुछ situation में models ने override यानि की निगरानी को कमजोर करने की कोशिश की। कुछ ने अपने किये गए
02:14action के बारे में बाद में सही जानकारी नहीं दी।
02:17तो कुछ test में model ने अपने weight यानि की model की copy को दूसरी जगा पर ले जाने की
02:22कोशिश की।
02:23यानि की model की copy पहले जहां थी वहां से चुरा करके कहीं और उससे paste करने की कोशिश की।
02:28ताकि बाद में वह अपने model को develop कर सके।
02:31सबसे दिल्चस्प बात थी कि कई बार model से सीथे तोर पर नहीं कहा गया था कि तुम cheating करो
02:36या तुम इंसानों को धोखा दो।
02:38situation ऐसी बनाई गई थी जिसमें ऐसा करना उसके दिये गए goal के लिए फायदे मन हो सकता था।
02:43यही चीज researcher के लिए इंट्रेस्टिंग थी। सवाल सिर्फ इतना नहीं था कि AI कोई गलत काम कर सकता है
02:49या नहीं।
02:50बलकि सवाल यह था कि क्या AI अपने goal को पूरा करने के लिए या जो उससे काम दिया गया
02:55है उस काम को पूरा करने के लिए
03:10Enthropic, Enthropic ने अपने cloud model के साथ भी ऐसे ही safety experiment किये। इन टेस्ट में researcher ने AI
03:16को एक simulated company environment दिया, AI को कुछ objective दिये गए और फिर ऐसी प्रिस्थितिया बनाई गई चहां model
03:23को लगा कि उसका role या access खत्म होने वाला है।
03:27कुछ scenario में model ने अपने objective को बचाने के लिए ऐसे action चुने जिनमें researcher ने हैरानी चता है।
03:57उसे change करने वाला था उसे blackmail करने की कोशिश की और बार बार धमकी दे रहा था। अब यहाँ
04:02पर थोड़ा ध्यान दिजी। यह असली company में हुआ कोई incident नहीं था।
04:06बलकि बास्टवीक एक्जेक्यूटिव AI के हाथो blackmail भी नहीं हुआ था। एक controlled test था। लेकिन researcher के लिए इसका
04:12महत्व इसलिए था क्योंकि model को सीधे instruction नहीं दिया गया था कि blackmail करो।
04:16situation ऐसी बनाए गई थी जहां model के goal को बचाने के लिए blackmail एक possible strategy बन रही थी।
04:22और यहीं से AI safety की पूरी debate और गंभीर हो जाती है।
04:26क्योंकि आज AI सिर्फ सवालों के जबाब देने वाला chatbot नहीं रह गया है। अब AI agent को emails पढ़ने,
04:33documents बनाने, software चलाने, code लिखने, website पर action लेने और अलग-अलग goals इस्तेमाल करने की capabilities भी दी
04:40जा रही है।
04:41मतलब AI को सिर्फ इनसान का दिमाग ही नहीं दिया जा रहा है, बलकि इनसान से बेहतर दिमाग दिया जा
04:47रहा है, और कुछ ऐसे दिमाग दिये जा रहे हैं, कुछ इस तरीके से विक्सित किया जा रहे हैं, कि
04:51जो काम इनसान काफी दिनों में कर सकता है, वो काम AI कुछी मिंट
04:55can be done.
04:57And when a system has more and more access, it can be more and more impact.
05:03Now, the question is, is that AI can be made by AI?
05:07No.
05:09This is a conclusion.
05:11Apollo Research and Anthropics test controlled environment.
05:15The experiment of how many capabilities and technologies can be done.
05:21और इसके साथ ही कि model अपने आपको बचाने के लिए क्या कुछ कर सकता है।
05:25तो इसका मतलब यह नहीं हुआ कि आप आज से समाने तोर पर एमान सके कि AI इंसानों के लिए
05:30खत्रा है लेकिन असली चिंता future की है।
05:32मान लिजिए आज AI के पास सिर्फ आपके सवाल का जबाब देनी की पावण है लेकिन कल वही AI किसी
05:37कमपनी के software सिस्टम चला रहा हो और उसके पास financial data हो।
05:42इंटर्नल इमेल सो हो, सरवर का एकसेस हो और उसे हजारो छोटे छोटे डिसेजन खुद लेने की permission है।
05:48अब अगर उस AI का goal गलत तरीके से सेट हो गया या किसी ने instruction को गलत समझ लिया
05:53तो problem बहुत बड़ी हो सकती है।
06:26अब सोचिए अगर कोई student exam में समझ ले कि teacher उसकी cheating पकड़ने वाला है तो आपना behavior बदल
06:33सकता है।
06:34AI के मामले में researcher भी इसी तरह की behavior को बहुत ध्यान से study कर रहे हैं।
06:38क्योंकि अगर model सिर्फ test के दौरान अच्छा behavior दिखाए और उसकी development में अलग behavior दिखाई दे यानि कि
06:45किसी अलग जगा पर उसका अलग behavior दिखाई दे तो safety testing मुश्कील हो सकती है।
06:50पोलर ने 2025 में open AI के साथ मिलकर frontier models पर anti-scaming training के test की। इन test
06:57में अलग-अलग models को controlled scenario में evaluate किया गया और उनके जूट बोलने की चीज़े बाकी चीज़ों और
07:04reward hacking जैसे behavior के example देखेंगे।
07:07लेकिन यहां भी वही बात लागू होती है कि यह control evaluation में थे रोजमर्रा की दुनिया में AI के
07:12समान behavior का सीधा प्रमान नहीं है। अब एक आम आदमी के लिए इसका क्या मतलब है।
07:16तो मतलब यह नहीं कि आपको आज से AI से ड़ना शुलू कर देना चाहिए बलकि मतलब यह है कि
07:21AI जितना powerful होता जा रहा है उतना ही जरूरी है कि उसके उपर control भी हो।
07:26जैसे की बैंक में करोरो रुपए का transaction किसी एक वेक्ती को बिना verification करने की इजाज़त नहीं दी जाती
07:32वैसे ही बहुत powerful AI को भी बिना monitoring के critical system का unlimited access देना risk पैदा कर सकता
07:40है।
07:41इसलिए researcher access control, monitoring, independent testing और shutdown mechanism जैसी चीजों पर चोड़ दे रहे हैं।
07:48और शायद आने वाले समय में AI की सबसे बड़ी completion सिर्फ यह नहीं होगी क्यों कौन सा model ज्यादा
07:53intelligent है।
07:54completion यह भी होगा कि कौन सा AI model ज्यादा reliable और controlled है।
07:59क्योंकि intelligence अकेले काफी नहीं है।
08:01अगर AI बहुत intelligent है लेकिन इनसान यह predict नहीं कर सकता कि वह किसी difficult situation में क्या करेगा
08:07तो problem पैदा हो सकती है।
08:09और यहीं Apollo और Enthropic की research का सबसे वरा take way है।
08:13AI ने अभी इनसानों के खिलाफ कोई स्वतंत्र युद्ध शुरू नहीं किया है।
08:17लेकिन control experiment ने दिखाया कि कुछ advanced model इन प्रेस्टिटियों में अपना अलग behavior देखा सकते है।
08:23तो इसलिए अगर आप भी AI का यूज करते हैं तो थोड़ा सा सतर्क रहिएगा और पोसिस किरिएगा कि AI
08:27के हाथों में पूरा control ना दे।
08:29और आपको वीडियो कैसा लगा कमेंट सेक्शन में जरूर बताईएगा और ऐसे ही तमाम अपडेट्स और शेयर मार्केट्स से जुड़ी
08:34तमाम जानकारी के लिए आप जुड़े रहें good returns दिश्टल के साथ

Recommended